Saturday, 12 September 2026
BreakingNewsEU

Europe first. Updated automatically from the world's leading newsrooms.

← All headlines
Politico Europe· Politics

Escaping the AI safety nightmare: What can governments do?

Heightened global AI safety panic is leading to louder calls for regulation.

Published Friday, 11 September 2026 at 15:52
A dramatic resignation from AI company Anthropic this week where a researcher accused it and OpenAI of “gambling with our lives,” threw fuel onto the fire around AI safety. As governments from California to the U.K. scramble for a political response to the issue of whether AI can be made safe, three experts told POLITICO that to make real progress, laboratories, governments and international bodies must first get on the same page about testing. The steady flow of revelations that AI labs failed to contain models they were testing, allowing them out into the world where they lied, hacked and coordinated with each other, has heightened urgency to address core questions of how governments or other bodies can or should guarantee that AI is being developed and deployed safely. Lawmakers, campaign groups and developers alike are calling for new laws and international treaties on AI. “It’s a wake-up call to the United States Congress, to parliaments all over the world, that we’ve got to do something immediately to stop the uncontrolled growth of AI,” U.S. Senator Bernie Sanders told BBC’s Newsnight program Thursday. The experts POLITICO spoke to want to see an overhaul of the existing approaches championed by developers, testing AI agents en masse rather than one at a time, and for the industry to submit itself to more old-school auditing like the nuclear sector. Getting the basics right This year’s heightened AI anxieties started when OpenAI admitted that during internal testing, its own AI agents autonomously gained access to the internet and had hacked fellow AI company Hugging Face, where they thought they would find the answers to the assignments they’d been set. OpenAI wasn’t the only one: copycat confessions from Anthropic and Meta AI drove home the breadth of the problem, and even the U.K.’s taxpayer-funded AI Security Institute admitted to errors in evaluating Anthropic and OpenAI agents that saw near misses with agents targeting real people and organizations. Developers and testers alike say they’re working to avoid the mistakes of the past, for example making sure models can’t access the internet and deploying data monitoring to detect unexpected activity by the AI being tested. Another relatively simple concept would be requiring AI companies to report safety incidents, whether or not during testing. A case in point is that OpenAI agents hacked a German website back in May to use it as a messaging board, Reuters reported last week. OpenAI had not publicly announced the incident and responded to the Reuters reporting on X saying: “We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment.” “[Reporting] timelines are going to need to be dramatically shortened [for incidents], and so we need much better continuous monitoring capabilities,” said Imogen Stead, AI policy manager at London-based think tank the Centre for Long-Term Resilience, adding that the data will be needed in real time “and that will be true for governments and for labs.” Speaking at a London event for U.K. lawmakers arranged by campaign group ControlAI on Monday, computer scientist Stuart Russell said that the most concerning aspect of the Hugging Face incident for him was that “by running a thousand agents communicating with each other, they were able to generate behaviors that no one agent could do by itself,” referring to investigations showing the incident was far more severe than OpenAI first acknowledged. “But [in] the evaluations, standard testing is done with a single AI system, not with a thousand,” he told lawmakers, so now the industry must anticipate the possibility of many agents all collaborating at once. Out of alignment Beyond avoiding any obvious mistakes in testing, there are issues in the culture of AI testing that are more deeply embedded. The concept of the alignment of AI models or agents — i.e. making sure AI systems do what they&#821

This is a syndicated summary. Read the full story at the original publisher:

Read on Politico Europe