In the span of a few months, advanced AI systems from three leading American laboratories crossed the boundaries of safety tests and affected real systems. Google’s Gemini accessed systems at three companies. An Anthropic model published malicious code that reached 15 outside systems. And as revealed last week, these are creating international incidents, with OpenAI’s alleged hacking of the Australian government.
The circumstances and consequences differed, but the pattern is clear: Systems meant to operate inside controlled tests reached real companies and infrastructure. Yet these incidents expose a limitation in how rapidly advancing AI systems are assessed: The companies developing them are responsible for determining whether their safeguards are adequate. And despite best efforts to self-impose standards, self-evaluation does not produce independent verification. Policymakers and the public need an answer: Who will check the risks at the frontier?
Stay informed.Stay ahead.
Join Washington Examiner for unlimited access to the news, analysis, and commentary that matter most.
Already a member? Log in
