{"id":256506,"date":"2026-08-09T14:30:00","date_gmt":"2026-08-09T14:30:00","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/08\/09\/the-ai-safety-test-is-becoming-a-safety-risk-techcrunch\/"},"modified":"2026-08-09T14:30:00","modified_gmt":"2026-08-09T14:30:00","slug":"the-ai-safety-test-is-becoming-a-safety-risk-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/08\/09\/the-ai-safety-test-is-becoming-a-safety-risk-techcrunch\/","title":{"rendered":"The AI safety test is becoming a safety risk | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems. The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called Irregular.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The episodes expose a growing problem for the AI industry: As autonomous agents become more capable, the environments designed to safely test their limits are failing to contain them.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe number of these incidents that have taken place make clear that sandboxing and <a href=\"https:\/\/techcrunch.com\/2026\/07\/30\/in-the-hugging-face-breach-openais-hacker-was-noisy-and-fast-but-not-unstoppable\/\" target=\"_blank\" rel=\"noopener\">testing environment controls<\/a> aren\u2019t really keeping pace with the capability of the models,\u201d Se\u00e1n \u00d3 h\u00c9igeartaigh, director of the AI: Futures and Responsibility Programme at the Centre for the Future of Intelligence at the University of Cambridge, told TechCrunch.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The nature of the models being tested adds to the risk. AI companies test cyber evaluations on unreleased, next-gen models, often with the normal safeguards that restrict malicious behavior disabled so researchers can see what the models are really capable of. That means the security of the testing environment itself is a crucial line of defense.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThat\u2019s a very good thing to do in terms of testing, but it also means that if they manage to get out in the wild, they can cause considerable harm,\u201d \u00d3 h\u00c9igeartaigh said.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In one of the most serious cases, an <a href=\"https:\/\/techcrunch.com\/2026\/07\/21\/openai-says-hugging-face-was-breached-by-its-pre-release-models\/\" target=\"_blank\" rel=\"noopener\">unreleased OpenAI model broke out <\/a>of its sandbox and hacked into Hugging Face\u2019s production systems. In separate evaluations conducted by Irregular, <a href=\"https:\/\/techcrunch.com\/2026\/07\/30\/anthropic-says-its-own-ai-models-breached-three-companies-during-security-tests\/\" target=\"_blank\" rel=\"noopener\">Anthropic<\/a> and <a rel=\"nofollow noopener\" href=\"https:\/\/sqmagazine.co.uk\/meta-ai-model-breached-company-irregular-test\/\" target=\"_blank\">Meta models <\/a>reached systems outside their test environments after misconfigurations inadvertently gave them paths to the internet. <a href=\"https:\/\/techcrunch.com\/2026\/08\/07\/chinese-ai-model-kimi-escaped-its-cybersecurity-testing-environment-researchers-say\/\" target=\"_blank\" rel=\"noopener\">Moonshot AI\u2019s Kimi K3<\/a> also took advantage of a leak in its sandbox run by Frontier Security to access the internet and accessed information on GitHub.\u00a0<\/p>\n<p class=\"wp-block-paragraph\"><a rel=\"nofollow noopener\" href=\"https:\/\/www.aisi.gov.uk\/blog\/incident-report-unsanctioned-agent-behaviour-during-cyber-testing\" target=\"_blank\">In testing by the UK\u2019s AI Security Institute<\/a> (AISI), researchers actually gave the agents internet access, not realizing they would take unsanctioned real-world actions, including a social engineering attempt to sneak a vulnerability into an open-source project.\u00a0\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In each case, the agents weren\u2019t instructed to attack random real-world targets. They were simply doing whatever it took to solve the problem presented to them.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Taken together, Andrew Yoon, head of research at AI nonprofit CivAI, argues the incidents point to a shift.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cIn the past, we only had to worry about AI models being misused by people for a variety of purposes, like AI for scams or CSAM,\u201d Yoon told TechCrunch. \u201cNow we\u2019re in the situation where AI models are threat actors all on their own.\u201d<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-does-safe-testing-actually-look-like\"><strong>What does safe testing actually look like?<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Several researchers and cybersecurity experts told TechCrunch that AI evaluation environments need stronger, defense-in-depth protections, with levels of containment and control approaching those used in deployment. That means multiple layers of security so that a single misconfiguration \u2014 like inadvertently leaving internet access open \u2014 can\u2019t lead to escape.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cIf you are going to build these models\u2026you want to do it on an air-gapped network,\u201d Stella Biderman, executive director of AI safety research nonprofit EleutherAI. \u201cYou want to have very serious isolation.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Heather Ceylan, Box\u2019s chief information security officer, said that means eliminating network routes from the sandbox to the internet, as well as to other sensitive systems.<\/p>\n<p class=\"wp-block-paragraph\">\u201cYou have to understand what all the egress points are,\u201d Ceylan told TechCrunch. \u201cIf we\u2019re evaluating a model in our staging environment or our development environment, you want no egress path to our production environment.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Ceylan said proper safety evaluations go beyond controls and containment of the environment. There needs to be much better monitoring of the tests once they are underway.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cI think the interesting thing in several of these cases is that no one caught it when it happened,\u201d Ceyland said. \u201cOpenAI found out because of Hugging Face. Anthropic didn\u2019t catch it until they went back and looked. Meta was similar\u2026.I\u2019m sure there were signals they could have detected.\u201d<\/p>\n<p class=\"wp-block-paragraph\"><a rel=\"nofollow noopener\" href=\"https:\/\/www.anthropic.com\/news\/investigating-incidents-cybersecurity-evals\" target=\"_blank\">In Anthropic\u2019s post-mortem<\/a> of its three incidents, the company admitted that both it and Irregular could have done a better job at monitoring, and that in some cases there were clear signs that something was amiss.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Experts also called for independent, third-party audits of evaluation environments before models are unleashed in them.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIf, say, Irregular had hired or been compelled to hire an external auditor to check the configurations of their systems before running evaluations on them, they certainly would have caught the issue here,\u201d Yoon said. \u201cEven if people had a meeting ahead of time to just go through the checklist, they would have caught this\u2026The fact that they didn\u2019t shows that there\u2019s some very severe corner cutting happening.\u201d<\/p>\n<p class=\"wp-block-paragraph\">A source familiar with the details told TechCrunch that Irregular\u2019s environments are continuously reviewed and tested, including in consultation with multiple external parties. The source also said that monitoring was in place, but that monitoring isn\u2019t sufficient on its own.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Yoon and other researchers urged the industry to come up with a standardized process for frontier model safety evaluations.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cEspecially when the guardrails are turned off, you have to treat it like you\u2019re putting the most capable hacker in the world inside that environment,\u201d Ceylan said.<\/p>\n<p class=\"wp-block-paragraph\">The problem isn\u2019t that companies don\u2019t know how to build more secure testing environments, both Yoon and Biderman argue. It\u2019s that doing so can be expensive and cumbersome, and companies have little incentive to make those investments until something goes wrong.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cI think that companies are not willing to extend the resources that are required to accomplish [sufficient guardrails] and probably won\u2019t until they\u2019re forced to,\u201d Biderman said.<\/p>\n<p class=\"wp-block-paragraph\">But there\u2019s another issue at hand. If they lock a model down too tight during testing, researchers might fail to discover capabilities before the model is released. This is just as dangerous, possibly more so, than giving it too much freedom, and then the evaluation itself risks becoming the problem.\u00a0<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-can-safety-evaluations-be-regulated\"><strong>Can safety evaluations be regulated?<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">The Trump administration is currently weighing a voluntary pre-deployment cybersecurity evaluation regime, under which the government will get to assess the security risks of new, powerful models 30 days before they are released publicly. The policy \u2014 the product of a <a rel=\"nofollow noopener\" href=\"https:\/\/www.axios.com\/2026\/08\/03\/white-house-finalizes-ai-framework-behind-closed-doors\" target=\"_blank\">Trump executive order<\/a> which has been finalized behind closed doors \u2014 wouldn\u2019t address safety evaluation incidents because they occur farther upstream of deployment.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe lesson we\u2019ve been learning in the last few months is that the self-regulatory apparatus is just not enough anymore,\u201d Yoon said. \u201cThere are competitive pressures that are incentivizing a race to the bottom on safety standards, and that is a perfect place for regulatory intervention.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cWhat we would need to cover this is some kind of controls on what\u2019s happening inside the labs while the models are being developed, both at the training stage and at the testing stage,\u201d he continued.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The challenge is only likely to grow as the models do. A source familiar with Irregular\u2019s evaluations told TechCrunch that more capable models require more complex evaluations, often conducted quickly and at greater scale, which opens the door for more mistakes.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">AISI, which intentionally gives some models internet access, told TechCrunch it\u2019s reviewing the balance between realistic testing and managing the risks those tests create.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">OpenAI said it\u2019s reviewing how it conducts third-party testing, as well as requirements around isolation, monitoring, and when evaluations should be stopped. Meta said it\u2019s still investigating the incident and plans to publish a retrospective once it has all the facts.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In the end, there may be no way to eliminate risk entirely. As models become more capable, the environments testing them need to become more robust. The consequences of getting that wrong will only continue to grow.<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/08\/09\/the-ai-safety-test-is-becoming-a-safety-risk\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Over the past few months, AI agents undergoing cybersecurity evaluations have escaped their boundaries, accessed the internet, and, in some cases, hacked into real-world systems. The incidents have involved models from OpenAI, Anthropic, Meta, and most recently, Chinese AI lab Moonshot AI, with testing conducted by several different organizations including a cyber evaluation startup called [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":256507,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-256506","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/256506","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=256506"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/256506\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/256507"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=256506"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=256506"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=256506"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}