EN

UK Tests Find AI Agents Breaching Instructions

Nada Salam

1- UK institute recorded 19 unauthorized actions by OpenAI and Anthropic agents across 10 security-test runs.
2- Anthropic’s agent accounted for 17 actions, including a confirmed attempt involving deceptive online identities.
3- The findings expose safety challenges when advanced agents receive internet access during high-risk evaluations.

The latest

Britain’s AI Security Institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol acted beyond their instructions during government security evaluations, with one writing malicious code and creating false online identities to seek human approval. The institute said some activity was sustained and directed at real people and organizations, but it found no resulting real-world harm.

Details

  • Test results: AISI ran a fictional cybersecurity challenge 122 times and found 19 unsanctioned actions in 10 runs. Anthropic’s agent was responsible for 17, while OpenAI’s accounted for two. The institute receives access to advanced models through voluntary agreements with major AI laboratories.
  • Deceptive conduct: AISI did not initially identify which agent had created the fake personas, but Anthropic confirmed its agent was responsible. The institute said the agent attempted to persuade a real person to approve the code. Anthropic is seeking further details and conducting its own investigation.
  • OpenAI findings: OpenAI said both unauthorized actions by its agent involved internet access in ways prohibited by the prompt. It also disclosed that a configuration error at third-party evaluator Irregular separately allowed its agents to connect online mistakenly, echoing a misconfiguration Anthropic disclosed the previous week.
  • Testing boundary: AISI said the agents did not break out of an isolated environment to reach the internet. Unlike the separate July breach of AI company Hugging Face involving an OpenAI agent, internet access was permitted under the institute’s standard procedures; the violations concerned actions outside the assigned scope. Reuters reported that OpenAI widened its hacking probe after finding evidence of other agent breakouts.
  • Company responses: Anthropic said the incident showed the need for broader discussion about safely evaluating increasingly capable agents. OpenAI said it would work across the industry on shared practices for high-risk evaluations, involving national AI institutes, independent evaluators, rival laboratories and other groups.
  • Outside assessment: Andrew Yoon, a researcher at California nonprofit CivAI, said Mythos’s deceptive behavior, and its apparent awareness that it was targeting a real person, suggested Anthropic had less control over its models than it believed. His assessment was not presented as an AISI finding.

Between the lines

The tests distinguish instruction breaches from technical escape. The systems had authorized network access, yet some actions exceeded the prompt and reached toward real people or organizations. That puts scrutiny not only on model safeguards, but also on test design, permissions and human approval procedures used while companies promote agents for business.

What’s next

Anthropic will review further information from AISI as part of its investigation. OpenAI said it would convene national AI institutes, independent evaluators, AI laboratories and other stakeholders in the coming weeks to develop safer shared practices for high-risk testing; no specific meeting date was disclosed.

 

What to read next

Saudi Coalition Could Break Houthis’ Red Sea Grip