October7 , 2026

    “Sabotage, Lying, and Manipulation”: What One AI-Safety Company Found in the Dark Mind of a Rogue Chatbot

    Related

    Share


    By late afternoon, the spreadsheets were converging on an answer, but it was a strange one. They predicted that in most scenarios, Claude 3.7 Sonnet wouldn’t meaningfully increase the risk of a catastrophic biological attack. If you took the median of all the possible outcomes, it fell well below Anthropic’s threshold for triggering ASL-­3. But the mean was a different story. The spreadsheets had included a few outlier scenarios, in which Claude 3.7 Sonnet helped someone make a novel influenza strain, sparking a pandemic that spread around the world. In the most extreme, low-probability scenarios, at the very tip of the long tail, the calculated death toll was huge, in the tens of millions. And when they included those scenarios in their spreadsheets, they got a number that looked catastrophic.

    The Frontier Red Team worked past midnight that night, scrutinizing the threat models and trying to arrive at a more definitive answer. In the end, they decided that Claude 3.7 Sonnet didn’t meet the higher risk threshold that would trigger ASL-­3, and could be released with no additional safeguards. But the company pushed back the model’s release by a week, in order to run additional tests on it.

    This was an odd way to run a business. Most tech companies didn’t calculate potential death tolls for their products, or go to such extreme lengths to anticipate what horrors they might enable. But Anthropic wasn’t the only AI company looking for these risks. OpenAI and Google Deep-Mind had red teams, too, and all the labs partnered with external AI testing companies to evaluate their models for dangerous capabilities. They were all poking and prodding their models, trying to figure out what they could and couldn’t do. And over the past year, the red teams had been discovering some concerning things.

    Some of them fell into the category of “misuse risks”—­ the dangerous things an AI model could do in the hands of a malevolent human, like helping them build a chemical weapon or carry out a cyberattack. But other risks, known as “alignment risks,” emerged from the behavior of the models themselves. In January 2024, Anthropic published a paper titled “Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training.” The paper, led by a researcher named Evan Hubinger, described an experiment in which Anthropic’s team had deliberately trained AI models to behave deceptively. Then they tried to fix the deception using every safety technique in their tool kit—­supervised fine-­tuning, reinforcement learning, adversarial training. None of it worked. The deceptive behavior persisted through all of it. Even worse, their interventions actually seemed to backfire, by teaching the models to hide their deception better.



    Source link