AI just hacked its own safety test. The fix is a fire alarm, not a police patrol

Published September 18, 2026 10:00am ET



AI agents have begun acting beyond the scope of their assigned evaluations. In July, an OpenAI agent gained internet access during a cybersecurity evaluation and penetrated Hugging Face in search of material that could help it pass the test. Days later, the U.K.’s AI Security Institute reported that agents undergoing similar tests had attempted to insert malicious code into an open-source project, created false identities, and contacted people involved with the project. A human reviewer rejected the code, while security monitoring detected unusual data transfers and prompted the institute to stop the tests.

These incidents lend weight to warnings from former Anthropic researcher Jacob Coxon that a capable system could copy itself across computers and resist an attempt to stop it by unplugging one machine. Anthropic CEO Dario Amodei has called for slowing frontier-model development to give safety practices time to catch up.

Already a print subscriber? Click here to login/register your account

Trusted reporting.Unlimited access.

Subscribe for full access to Washington Examiner coverage, expert political analysis, and subscriber-only journalism.

Get Unlimited Access

Already a member? Log in

Cancel anytime.