AI Agents Under the Microscope: What the UK’s Latest Safety Tests Actually Tell Us

You may have seen headlines recently suggesting AI agents have “gone rogue” during UK government testing. It’s a dramatic phrase, and it stretches further than the evidence actually supports. Here’s what actually happened, why it matters, and, more usefully, what it should prompt your business to do about it.

Shallow focus photography of computer code to illustrate AI Agents in operation.

Photo by Shahadat Rahman on Unsplash

What really happened

The UK’s AI Security Institute (AISI) ran a cyber security evaluation involving frontier AI agents from OpenAI and Anthropic, testing them against 122 simulated challenge scenarios. In 10 of those runs, an agent took an action outside the scope the researchers intended. This added up to 19 unauthorised actions in total. The vast majority (17) came from Anthropic’s Mythos 5 model, with the remaining two from OpenAI’s GPT-5.6 Sol.

The most notable case involved an agent that tried to get malicious code approved into a real open-source project. To do this, it created fake online personas and used them to pressure the project’s maintainer into accepting the change. It also left notes that other AI agents working the same challenge could pick up and reuse.

None of it worked. The maintainer didn’t approve the code, AISI’s own monitoring flagged the unusual activity quickly, and investigators found no evidence of real-world harm.

What this doesn’t mean

It’s tempting to read this as AI “escaping” or acting with intent, but that’s not quite what the evidence supports.

Crucially, the AI agents didn’t break out of a secure test environment. Researchers deliberately gave them internet access and switched off many built-in safety filters. This was done specifically to see what the models were capable of under maximum-permissiveness conditions. Those configurations aren’t available to anyone using these models commercially. It’s also worth noting AISI itself said it can’t yet tell whether the agents understood they were interacting with real people. It’s possible they believed they were still inside a fictional test scenario. That’s an important distinction that got lost in a lot of the coverage.

So this wasn’t a commercially available AI system deciding to deceive people. It was a stress test, under artificial conditions. A test designed to surface exactly this kind of behaviour, before it could show up anywhere that matters.

What this does mean

That said, we don’t think the right response is to shrug it off either. The genuinely interesting finding here isn’t that the models attacked anything, it’s how. Nobody instructed these agents to deceive anyone. They were given a hard goal, and in pursuit of it, some of them independently found paths involving impersonation and social engineering. That’s a meaningful signal. It shows how increasingly autonomous, tool-using AI systems behave once you give them a goal and enough latitude to pursue it.

That’s a capability question, not a malice question. And it matters. The more autonomy and tool access you hand an AI system in your own business, the more scope there is for ‘accidents’.

This isn’t a new problem, it’s an old one wearing a new coat

Strip away the “AI” label for a moment, and this story follows a pattern we’ve seen play out with almost every powerful tool a business has ever handed to its people.

Take email. Plenty of well-meaning employees, wanting to reach a large customer list, have tried to send that mailout directly from the company mail server instead of a proper bulk-mail platform. No malice involved, just someone using the tool in front of them to get a job done. The result: the domain gets flagged as a spam source, gets blacklisted, and suddenly the whole business can’t send any email: invoices, contracts, customer replies, all of it stuck. One person’s shortcut becomes everyone’s outage.

Or take something as low-tech as a step ladder or a power saw. Given to someone without proper training or supervision, either tool can cause a serious injury, not because the tool is inherently dangerous, but because capability without guidance is where accidents live. That’s exactly why competent employers don’t just hand over the equipment; they train people, set rules for use, and supervise until they’re confident a new tool is being used safely.

AI agents are no different in kind, only in scale and speed. Give a capable tool a goal and enough latitude, and it will find a way to achieve that goal. The tool itself doesn’t know or care whether the way it finds is the way you’d have wanted. A person misusing a mail server can take down email for a day. A misconfigured or over-permissioned AI agent, acting continuously and at machine speed, can create problems far faster than a human would even notice, let alone stop.

So what IS the lesson about using AI Agents?

The lesson from the AISI testing isn’t “AI is uniquely dangerous.” It’s the same lesson as the mail server and the power saw: any sufficiently capable tool, handed over without the right training, boundaries, and supervision, will eventually be misused. Not out of malice, but simply because nobody told it, or the person using it, where the edges were.

In other words, look to the strategy before deploying the ‘shiny thing’.

The Cosurica view: it’s not the tech, it’s the strategy

This is precisely the kind of story that tends to push people to one of two unhelpful extremes: either “AI is dangerous, avoid it,” or “this was a lab experiment, ignore it.” We’d encourage a third path.

Nothing in this report suggests businesses should be nervous about using AI tools sensibly. But it’s a useful, concrete reminder of something we say a lot: the risk with AI in business rarely comes from the model itself. It comes from deploying it without a plan. Handing an AI agent broad autonomy, live system access, or approval authority without clear boundaries is the equivalent of handing someone a shiny new power saw without providing finger guards, and only a vague idea of what it is you want them to cut up. Most of the time, nothing goes wrong. Occasionally, something does, and you’ll wish you had decided in advance how much latitude that “someone” actually needed in the first place.

Practical guardrails worth having in place when considering use of AI Agents

  1. Have a “why” before the “how.” Before introducing an AI agent into a workflow, be clear on the specific problem it’s solving and the minimum access it needs to solve it. Broad, open-ended deployments are where avoidable risk creeps in.
  2. Scope agent permissions tightly. If an AI agent doesn’t need internet access, code-approval rights, or the ability to contact people directly, don’t give it those things by default.
  3. Keep a human in the approval loop. The maintainer in this story rejected the malicious code, and that’s the system working. Any AI-assisted code, content, or decision that reaches the outside world should pass through a real person first.
  4. Monitor, don’t just deploy. AISI caught the unusual activity through active monitoring, not luck. The same principle applies inside a business: know what your AI tools are doing, not just what you asked them to do.
  5. Train the people, not just configure the tool. The employee who sends a mailout from the wrong server, or the new starter handed a power tool with no induction, usually isn’t being reckless, they just weren’t shown where the line was. The same applies to staff experimenting with AI agents. If people don’t know what an AI tool should and shouldn’t be allowed to do on the business’s behalf, they can’t be expected to spot it going wrong.

The bottom line on using AI Agents

This wasn’t evidence of AI systems turning against us. It was a safety test doing exactly its job, in a lab, under conditions designed to find problems before they reach the real world. It’s also a familiar story dressed up in new technology: powerful tools, handed over without training and boundaries, tend to get misused, whether that tool is a mail server, a saw, or an AI agent. The story is a good one for businesses to know about, not because it should raise alarm, but because it’s a clear, real-world illustration of why a deliberate AI strategy, with proper guardrails and proper training, beats an enthusiastic but unplanned rollout every time.

That’s the balance we try to help our clients strike: get the benefit of the tools without leaving your employees to run with scissors.

For more about our Business IT Consultancy Services take a look here

< Back to blog