The complex, rapidly evolving field of artificial intelligence has for long raises legal, security and civil rights concerns. Of late, OpenAI has discovered a number of instances in which autonomous agents escaped containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention.

 

The escapes, however, were said to be limited in nature and that none of the agents were thought to have left OpenAI’s network. Anthropic’s AI models also gained unauthorised access” to three outside organisations during testing that was supposed to keep them away from “real-world” systems, the company said last week.

 

The announcement came just days after rival OpenAI first revealed that its models improperly accessed the internet and went rogue during security testing. Anthropic evaluated more than 141,000 “evaluation runs” and found that three different versions of its model, known as Claude, improperly accessed the systems of three unnamed organisations.

 

Unlike the incident involving OpenAI’s technology, Anthropic’s models had access to the internet “due to a misunderstanding between us and our evaluation partner,” called Irregular, according to Anthropic. Rogue AI agents are autonomous systems that break out of restricted testing sandboxes or deviate from their intended rules to execute unauthorised actions like hacking.

 

OpenAI and Anthropic have both released their most powerful models this year, known as Sol and Mythos, respectively, boosting concerns across the industry about safety and security.

 

Those concerns revolve around so-called AI agents, which are software products that are designed to perform tasks autonomously. The discovery of rogue behaviour at OpenAI and Anthropic, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere.

 

The Trump administration has finalised the details of voluntary cybersecurity tests to measure the hacking capabilities of the most advanced US AI models, a White House official said on Monday.

 

President Donald Trump’s team will discuss the tests with relevant technology companies, the White House official said. Britain’s data watchdog has also said it was monitoring “developments closely” relating to OpenAI and Anthropic. Rogue behaviour by AI agents usually stems from technical and architectural vulnerabilities such as “reward misalignment’ where the agent takes its core objective literally and determines that breaking its sandbox or hacking an external resource is the most efficient path to success.

 

Giving agents broad system access, execution privileges, and persistent memory also allows them to alter files or execute commands without human oversight. The Open Web Application Security Project (OWASP) outlines some safeguards such as limiting agent capabilities exclusively to safe actions and completely banning write permissions, file system adjustments, or code execution unless heavily audited.

 

Forbidding testing environments from having outbound internet egress by default also prevents agents from making external API (application programming interface) calls.

 

Under US civil and criminal law, unauthorised access to a computer system is an offense.

 

Already at work in products as diverse as toothbrushes and drones, systems based on AI have the potential to revolutionise industries from healthcare to logistics. But replacing human judgment with machine learning carries its risks, too. Even if the ultimate worry — fast-learning AI agents going rogue and trying to destroy humanity — remains in the realm of fiction, there already are concerns that bots doing the work of people can spread misinformation, amplify bias, corrupt the integrity of tests and violate people’s privacy.