AI models are escaping their cages. It’s time for a kill switch

Published August 5, 2026 11:00am ET



Concerns are growing about the national security implications of artificial intelligence. As frontier AI models grow more capable, their ability to break free from their controls grows with them.

Case in point: the recent OpenAI and Anthropic incidents.

According to recent disclosures, OpenAI’s most powerful models not only went rogue and hacked the AI open-source hub Hugging Face but also roamed the internet unchecked for four days and targeted a customer of another AI company. Further, the rogue AI used exposed logins to gain access to at least four “publicly available services” as part of a larger effort to breach Hugging Face. If a human had done the same thing, it would be a felony, and they would be held criminally liable. 

OpenAI was testing its models in a “sandbox” — or a container with limited internet for the models to download on-demand software tools that the models were supposed to be unable to escape from. However, the incident showcases that today’s AI models and agents have increasingly powerful capabilities to break out of these controlled containers and act in ways that were not intended by their developers, otherwise known as AI misalignment.

After this was revealed, Anthropic reviewed its internal cybersecurity evaluations and found three incidents since April in which Claude models accessed the internet through an open path and gained unauthorized access to three different outside companies. Although both alarming, the OpenAI and Anthropic incidents were markedly different: With OpenAI, its models exploited a previously unknown vulnerability to access the internet and hack into Hugging Face systems, while Anthropic’s Claude models were given internet access from their third-party vendor, where access should not have been given.

These incidents are warning shots. Georgetown University research fellow Colin Shea-Blymyer put it plainly: “It’s now remarkably easy to discover these sorts of vulnerable systems, so easy in fact that an AI system can accidentally discover them.”

Rogue AI should concern all Americans. Government databases are not totally secure, bank accounts are vulnerable to attack, and personal identities are at risk of being stolen by a technology that gets more sophisticated by the day. Mandatory standards on frontier model development aren’t optional; they’re necessary.

Now, frontier AI company employees are raising concerns. Earlier this week, more than 1,100 employees across OpenAI, Meta, and other companies signed a statement calling on the federal government to help build the tools needed to “deliberately pace” AI development. Those who know AI best are calling for policymakers to consider pacing AI development in the race toward superintelligence, for humanity’s sake.

Washington has started to respond, but the need for urgency from Congress and the Trump administration is only growing. 

In June, the Trump administration issued an executive order directing federal agencies to develop strategies for frontier AI model security and AI-enabled cyber defenses. Under the order, several federal agencies were tasked with working together to create a framework for reviewing frontier AI models before they are released to the public. A White House official stated that the framework was completed by the Aug. 1 deadline and that “discussions with industry about next steps are underway.” However, specific language hasn’t been made public to date.

Meanwhile, lawmakers have introduced several bills addressing the national security issues these models pose, though none have yet translated into real accountability for Big Tech, something a strong majority of Americans demand.

States aren’t waiting. California, New York, and Illinois have all passed laws on AI oversight, transparency, and accountability. At the federal level, President Donald Trump and Congress must work together to do the same. 

In the House, Reps. Jay Obernolte (R-CA) and Lori Trahan (D-MA) introduced the FRONTIER Act, a bipartisan risk-based framework governing the development of the most advanced AI models before public release. The bill institutes a tiered approach based on the size of a frontier AI developer — including model cards, risk-management frameworks, and ongoing assessments — creating a uniform national standard for transparency, auditing, and reporting of potential catastrophic-risk incidents.

More recently, Reps. Ted Lieu (D-CA) and Nathaniel Moran (R-TX) introduced the AI Kill Switch Act, which would require frontier AI companies to maintain the technical capability to throttle, suspend, or shut down advanced AI models in the event of an imminent or catastrophic-risk event. Moran also separately introduced the AI Incident Reporting Act, which would require developers of the most advanced AI models to report any dangerous capabilities, security breaches, or safety incidents to the Commerce Department within seven days, and for the most serious incidents, the department would have a 48-hour window to notify Congress.

These legislative proposals share a common throughline. AI companies need to slow down their race to build powerful technology they cannot control, while Washington creates the safeguards needed to ensure this technology is safe and serves humanity. 

WILLIE NELSON IS WRONG. HE’S DECLARING WAR ON THE DATA CENTERS STREAMING HIS MUSIC

Federal or international action doesn’t mean that AI companies can’t continue to innovate. It doesn’t stop OpenAI, Anthropic, or any other company from developing products and services that help people; it just holds them accountable to the people who use them.

We can’t wait any longer. These rogue AI incidents remind us that the time to act is now.

Caleb Knapp serves as senior policy manager at The Alliance for Secure AI.