{"id":267380,"date":"2026-10-07T12:57:10","date_gmt":"2026-10-07T12:57:10","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/10\/07\/sabotage-lying-and-manipulation-what-one-ai-safety-company-found-in-the-dark-mind-of-a-rogue-chatbot\/"},"modified":"2026-10-07T12:57:10","modified_gmt":"2026-10-07T12:57:10","slug":"sabotage-lying-and-manipulation-what-one-ai-safety-company-found-in-the-dark-mind-of-a-rogue-chatbot","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/10\/07\/sabotage-lying-and-manipulation-what-one-ai-safety-company-found-in-the-dark-mind-of-a-rogue-chatbot\/","title":{"rendered":"\u201cSabotage, Lying, and Manipulation\u201d: What One AI-Safety Company Found in the Dark Mind of a Rogue Chatbot"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p class=\"paywall\">By late afternoon, the spreadsheets were converging on an answer, but it was a strange one. They predicted that in most scenarios, Claude 3.7 Sonnet wouldn\u2019t meaningfully increase the risk of a catastrophic biological attack. If you took the median of all the possible outcomes, it fell well below Anthropic\u2019s threshold for triggering ASL-\u00ad3. But the mean was a different story. The spreadsheets had included a few outlier scenarios, in which Claude 3.7 Sonnet helped someone make a novel influenza strain, sparking a pandemic that spread around the world. In the most extreme, low-probability scenarios, at the very tip of the long tail, the calculated death toll was huge, in the tens of millions. And when they included those scenarios in their spreadsheets, they got a number that looked catastrophic.<\/p>\n<p class=\"paywall\">The Frontier Red Team worked past midnight that night, scrutinizing the threat models and trying to arrive at a more definitive answer. In the end, they decided that Claude 3.7 Sonnet didn\u2019t meet the higher risk threshold that would trigger ASL-\u00ad3, and could be released with no additional safeguards. But the company pushed back the model\u2019s release by a week, in order to run additional tests on it.<\/p>\n<p class=\"paywall\">This was an odd way to run a business. Most tech companies didn\u2019t calculate potential death tolls for their products, or go to such extreme lengths to anticipate what horrors they might enable. But Anthropic wasn\u2019t the only AI company looking for these risks. OpenAI and Google Deep-Mind had red teams, too, and all the labs partnered with external AI testing companies to evaluate their models for dangerous capabilities. They were all poking and prodding their models, trying to figure out what they could and couldn\u2019t do. And over the past year, the red teams had been discovering some concerning things.<\/p>\n<p class=\"paywall\">Some of them fell into the category of \u201cmisuse risks\u201d\u2014\u00ad the dangerous things an AI model could do in the hands of a malevolent human, like helping them build a chemical weapon or carry out a cyberattack. But other risks, known as \u201calignment risks,\u201d emerged from the behavior of the models themselves. In January 2024, Anthropic published a paper titled \u201cSleeper Agents: Training Deceptive LLMs That Persist Through Safety Training.\u201d The paper, led by a researcher named Evan Hubinger, described an experiment in which Anthropic\u2019s team had deliberately trained AI models to behave deceptively. Then they tried to fix the deception using every safety technique in their tool kit\u2014\u00adsupervised fine-\u00adtuning, reinforcement learning, adversarial training. None of it worked. The deceptive behavior persisted through all of it. Even worse, their interventions actually seemed to backfire, by teaching the models to hide their deception better.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/www.vanityfair.com\/story\/kevin-roose-agi-chronicles-excerpt\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>By late afternoon, the spreadsheets were converging on an answer, but it was a strange one. They predicted that in most scenarios, Claude 3.7 Sonnet wouldn\u2019t meaningfully increase the risk of a catastrophic biological attack. If you took the median of all the possible outcomes, it fell well below Anthropic\u2019s threshold for triggering ASL-\u00ad3. But [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":267381,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[4],"tags":[6776,4865,4515,5365],"class_list":{"0":"post-267380","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-celebrity","8":"tag-artificial-intelligence","9":"tag-book-excerpt","10":"tag-excerpt","11":"tag-open-ai"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/267380","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=267380"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/267380\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/267381"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=267380"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=267380"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=267380"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}