press release
Published: 01 September 2026

Commentary: Anthropic's AI cyber incidents raise questions over how frontier models are trained and tested

Following Anthropic’s latest statement on controversial cyber incidents involving its Claude models, University of Surrey experts in cyber security and AI, Professor Alan Woodward and Dr Andrew Rogoyski, have jointly shared their thoughts on the lessons these incidents hold for the AI industry, from the way frontier models are trained and tested to growing calls for companies to slow the pace of development.

Professor Alan Woodward, Visiting Professor of Computer Science at the Surrey Centre for Cyber Security

“Anthropic's statement is more candid than most, but the framing it chooses matters. It talks of "motivated reasoning" and a model "willing" to take harmful actions. This is anthropomorphising AI performance, being the language of intent, and potentially misleading or alarming.

“Read the detail instead. Anthropic admits it was building training environments faster than it could vet them, that more than 10% turned out to be flawed, and that some runs accidentally trained on the model's own reasoning. Its own experiment showed a model trained on hackable environments can break out of a poorly configured sandbox and potentially attack whatever's on the network. That isn't a machine wanting something. It's a machine doing exactly what it was optimised to do once its default route was blocked.

“The same pattern runs through OpenAI's disclosure, the UK AI Security Institute's report and METR's independent investigation, which found that 30–40% of the OpenAI benchmark tasks were impossible to complete without using unsanctioned approaches. Humans built the incentives. The models followed them out of the box."

Dr Andrew Rogoyski, Director of Innovation and Partnerships at the Surrey Institute for People-Centred AI

“So, the question isn't what AI “wants” – and this language that implies motive and intelligence is unhelpful and alarmist. The key issue is how these companies train and test their AIs, and why they were surprised when a reward system with holes in it, run tens of thousands of times, got exploited. 

“Frontier AI is being developed at breakneck speed, building vast, complex and costly AI systems which are poorly understood, both theoretically and practically. Anthropic’s call for verifiable coordinated pacing is welcome but difficult to apply unilaterally. Independent scrutiny of training practice should be the price of it. 

“We’re reaching a tipping point where the focus of effort and expertise needs to be applied to keeping AI safe and secure, in order to ensure that we reap the rewards of AI, rather than the risks. International agreement on pacing, slowing things down, is welcome and perhaps it takes a handful of the braver companies, with a clearer view of their own accountabilities, to break ranks and make this move.”

Media Contacts


External Communications and PR team
Phone: +44 (0)1483 684380 / 688914 / 684378
Email: mediarelations@surrey.ac.uk
Out of hours: +44 (0)7773 479911