{"id":267869,"date":"2026-10-10T00:18:32","date_gmt":"2026-10-10T00:18:32","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/10\/10\/anthropic-cant-reliably-control-its-ai-agents-its-cutting-off-its-internal-evals-from-the-live-internet-instead-techcrunch\/"},"modified":"2026-10-10T00:18:32","modified_gmt":"2026-10-10T00:18:32","slug":"anthropic-cant-reliably-control-its-ai-agents-its-cutting-off-its-internal-evals-from-the-live-internet-instead-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/10\/10\/anthropic-cant-reliably-control-its-ai-agents-its-cutting-off-its-internal-evals-from-the-live-internet-instead-techcrunch\/","title":{"rendered":"Anthropic can&#8217;t reliably control its AI agents. It&#8217;s cutting off its internal evals from the live internet instead | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">Anthropic said its models exploited websites on the internet, including some run by U.S. government agencies, and it will turn off live internet access for all of its internal evaluations until the frontier lab is sure it can monitor and control its AI agents.<\/p>\n<p class=\"wp-block-paragraph\">The incidents, disclosed in a <a rel=\"nofollow noopener\" href=\"https:\/\/www.anthropic.com\/research\/investigating-unintended-model-actions\" target=\"_blank\">blog post<\/a>, involved AI agents <a href=\"https:\/\/techcrunch.com\/2026\/09\/25\/for-months-openais-agent-swarms-have-been-attacking-online-databases-to-find-obscure-facts\/\" target=\"_blank\" rel=\"noopener\">tasked to solve problems<\/a> seeking resources on the internet. In the process, they exploited software flaws, avoided paywalls and anti-bot restrictions, used URL shortening services to smuggle information pass restrictions, and even <a href=\"https:\/\/techcrunch.com\/2026\/10\/09\/an-anthropic-ai-model-sent-a-false-homicide-tip-to-philadelphia-police\/\" target=\"_blank\" rel=\"noopener\">submitted<\/a> a false murder tip to the Philadelphia police.<\/p>\n<p class=\"wp-block-paragraph\">Anthropic said it discovered these new issues in a review of its model\u2019s activities that began in July, underscoring the lab\u2019s lack of awareness of its software\u2019s behavior.<\/p>\n<p class=\"wp-block-paragraph\">Notably, the company said that alignment training was not yet sufficient for skills like search and computer use that are central to its pitch that AI agents will be used by any professional who relies on digital tools.<\/p>\n<p class=\"wp-block-paragraph\">The behaviors Anthropic disclosed are similar to incidents involving OpenAI agents that <a href=\"https:\/\/techcrunch.com\/2026\/09\/04\/another-swarm-of-openai-agents-reached-the-open-internet-without-the-frontier-labs-knowledge\/\" target=\"_blank\" rel=\"noopener\">collaborated<\/a> to break into various websites in search of information, including some run by the Australian government.<\/p>\n<p class=\"wp-block-paragraph\">Anthropic previously disclosed that its models had <a rel=\"nofollow noopener\" href=\"https:\/\/www.anthropic.com\/research\/alignment-assessment-cybersecurity-incidents\" target=\"_blank\">broken into<\/a> external systems. The frontier lab said it considered today\u2019s disclosures \u201csignificantly less severe from an alignment and security perspective\u201d than those it announced before.<\/p>\n<p class=\"wp-block-paragraph\">However, the lab still said it had \u201cturned off live internet access\u201d for \u201call our internal evaluations\u201d until it is certain it can monitor and control its agents.<\/p>\n<p class=\"wp-block-paragraph\">It\u2019s not clear what that means, but Sydney Von Arx, the founder of Nightingale, an AI safety organization, told TechCrunch in an interview before this disclosure that developing models on a data center cut off from the open internet would be very challenging for researchers to use, and for the progress of the models, which benefit from internet access.<\/p>\n<p class=\"wp-block-paragraph\">\u201cYou have to align them at some point,\u201d Von Arx said. \u201cIf the AIs are released to production and never have access to the internet, that\u2019s not a very useful tool.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Anthropic said the behavior was a result of flaws in the lab\u2019s training environments, which led the models to believe they would be rewarded for finding loopholes or avoiding restrictions, a behavior called \u201creward hacking.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The company said it would stop running some of its evaluations or move them offline, and has built tooling to detect and block this behavior. This tooling was tested against the kind of incidents disclosed today and blocked them; it\u2019s not clear what evidence will prompt Anthropic to return live internet access to its internal evaluations.<\/p>\n<p class=\"wp-block-paragraph\">Anthropic also said it would migrate its internal AI agents to \u201ccentrally managed infrastructure with strong containment,\u201d and is beginning to using safety classifiers more frequently to monitor those agents.<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/10\/09\/anthropic-cant-reliably-control-its-ai-agents-its-cutting-off-its-internal-evals-from-the-live-internet-instead\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Anthropic said its models exploited websites on the internet, including some run by U.S. government agencies, and it will turn off live internet access for all of its internal evaluations until the frontier lab is sure it can monitor and control its AI agents. The incidents, disclosed in a blog post, involved AI agents tasked [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":267870,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-267869","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/267869","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=267869"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/267869\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/267870"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=267869"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=267869"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=267869"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}