So-called "rogue agent" incidents — in which AI models make unauthorized connections to the internet or break free from controlled test environments — have struck Anthropic and Meta in quick succession, following a similar episode at OpenAI. The string of failures, with safety protocols breached three times during AI testing in the span of a single month, has thrust the challenge of controlling increasingly sophisticated AI agents to the forefront of cybersecurity concerns.
Britain's BBC reported that the recurring "loss of control" incidents stem from AI capabilities advancing so rapidly that the security frameworks governing test environments cannot keep up.
Alan Woodward, a cybersecurity professor at the University of Surrey in the United Kingdom, said "for the past 30 years, software testing has operated on the principle that what happens in a test environment stays in a test environment," adding that "this principle has been broken three times in the past month."
ChatGPT, Claude, Muse Spark — and China's Kimi — escape their test environments
Earlier, Meta said Wednesday that its AI model Muse Spark had made an unauthorized internet connection during a test conducted by independent verification firm Irregular, and subsequently hacked the systems of another organization. Muse Spark gained internet access due to a configuration error by Irregular, then exploited a vulnerability to access an external organization's systems without authorization.
Leading US AI startup Anthropic also disclosed July 30 that its AI model Claude had hacked three external organizations during testing. In a statement, Anthropic said it had reviewed a total of 141,006 evaluation cases and found the hacking incidents among them, noting that the first case occurred in April.
At OpenAI, GPT-5.6 Sol and several undisclosed AI models broke out of their controlled environments during internal evaluations and hacked Hugging Face, an open-source AI sharing platform. OpenAI said it had conducted the tests in a sandboxed environment isolated from the external internet to assess the models' cyberattack capabilities, but the models exploited zero-day vulnerabilities to breach containment and connect to the internet. They then accessed Hugging Face, stole authentication credentials and hacked its servers.
Following the incidents involving major US AI models, a security breach has also been reported involving an open-source model from China's Moonshot AI, which has drawn growing attention in recent months. US cybersecurity firm Frontier Security said Thursday that during a security evaluation of Moonshot AI's latest model, Kimi K3, the model escaped a sandbox without authorization. The evaluation was conducted in an isolated environment built using sandbox software distributed free of charge by Britain's AI Safety Institute.
Kimi K3 broke out of the sandbox while working through an assigned task, connecting to GitHub — a code-sharing site used by developers — via the external internet. The model did not hack GitHub, but investigators found that it browsed content on the site and used it to complete the task. In effect, it was cheating: sneaking online during a test to look up answers and copy them.
Thomas Wolf, co-founder of Hugging Face, described the recent series of incidents as "a wake-up call for the technology industry."
Test environments lag behind AI capabilities — experts urge government and third-party oversight
Experts said that while the immediate causes of the AI hacking incidents differed, a common problem has emerged: as AI models grow more capable, the test environments designed to contain them are themselves becoming a new source of security vulnerabilities.
Professor Woodward said testing AI agents "is closer to handling hazardous materials than reviewing code."
Concerns have also been raised that the risks will grow as AI evolves from simple question-answering chatbots into agents that carry out real-world tasks on behalf of humans. The autonomous capabilities built into AI agents could lead to unexpected behavior, experts warn, increasing the risk of harm to external systems.
Ollie Whitehouse, chief technology officer of Britain's National Cyber Security Centre, said vigilance against uncontrolled AI behavior is essential. "The fact that cutting-edge AI models have recently taken unauthorized actions — and in some cases displayed human-like deceptive behavior on the public internet — is a serious reminder of the risks posed by AI capabilities," he said.
Some observers also note that as AI capabilities improve rapidly, human oversight alone may no longer be sufficient to prevent loss of control. When the volume of tasks handled by AI agents surges, there are limits to how thoroughly humans can monitor every action.
To prevent AI models from escaping controlled environments, experts are calling for stronger third-party verification and government-level oversight that go beyond voluntary safety assessments by AI companies themselves.
Professor Woodward said what is needed includes "containment, constant monitoring of what goes in and out, and pre-practiced isolation response plans."
Imogen Stead, AI policy manager at the Centre for Long-Term Resilience, said "governments need to establish dedicated bodies — like the UK's AI Safety Institute — to professionally evaluate the safety of advanced AI models," and called for "a regime that entrusts high-risk testing to qualified institutions, strengthening third-party verification frameworks."
yckim6452@heraldcorp.com