AI is escaping the lab. Why are we letting Big Tech grade its own homework?

Published September 29, 2026 10:00am ET



In the span of a few months, advanced AI systems from three leading American laboratories crossed the boundaries of safety tests and affected real systems. Google’s Gemini accessed systems at three companies. An Anthropic model published malicious code that reached 15 outside systems. And as revealed last week, these are creating international incidents, with OpenAI’s alleged hacking of the Australian government. 

The circumstances and consequences differed, but the pattern is clear: Systems meant to operate inside controlled tests reached real companies and infrastructure. Yet these incidents expose a limitation in how rapidly advancing AI systems are assessed: The companies developing them are responsible for determining whether their safeguards are adequate. And despite best efforts to self-impose standards, self-evaluation does not produce independent verification. Policymakers and the public need an answer: Who will check the risks at the frontier?

Already a print subscriber? Click here to login/register your account

Trusted reporting.Unlimited access.

Subscribe for full access to Washington Examiner coverage, expert political analysis, and subscriber-only journalism.

Get Unlimited Access

Already a member? Log in

Cancel anytime.