{"id":263358,"date":"2026-09-16T18:25:25","date_gmt":"2026-09-16T18:25:25","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/16\/ai-labs-want-in-house-auditors-but-maybe-they-should-shut-the-front-door-first-techcrunch\/"},"modified":"2026-09-16T18:25:25","modified_gmt":"2026-09-16T18:25:25","slug":"ai-labs-want-in-house-auditors-but-maybe-they-should-shut-the-front-door-first-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/16\/ai-labs-want-in-house-auditors-but-maybe-they-should-shut-the-front-door-first-techcrunch\/","title":{"rendered":"AI labs want in-house auditors \u2014 but maybe they should shut the front door first | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei wrote about the need for outside organizations \u201cto verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.\u201d Executives at OpenAI, Google and SpaceXAI have already <a href=\"https:\/\/techcrunch.com\/2026\/09\/15\/openai-anthropic-google-have-been-in-talks-on-ai-safety-for-weeks\/\" target=\"_blank\" rel=\"noopener\">rallied around<\/a> Amodei\u2019s plan, which has quickly become a central pillar of the emerging AI safety push.<\/p>\n<p class=\"wp-block-paragraph\">But there may be a simpler and more effective fix hiding in plain sight. Internet security experts say the labs need to focus on network security basics like logs and permissions, applying the same rigorous defenses they do for human users. It\u2019s not as exciting as third-party auditing and alignment work\u2014but it may end up being more effective.<\/p>\n<p class=\"wp-block-paragraph\">\u201cTo me, it seems like they\u2019re outsourcing,\u201d Kate Moussoris, the CEO of Luta Security, told TechCrunch of Amodei\u2019s proposal. \u201cSaying [a third-party audit] is the solution is a strange proposition from my perspective. It would be the same as if, instead of writing the <a rel=\"nofollow noopener\" href=\"https:\/\/www.wired.com\/2002\/01\/bill-gates-trustworthy-computing\/\" target=\"_blank\">Trustworthy Computing Memo<\/a>, Microsoft said, let\u2019s slow down development.\u201d<\/p>\n<p class=\"wp-block-paragraph\">That memo, written by then-Microsoft CEO Bill Gates in 2002, called on his employees to ensure that their software would be reliable and safe following a series of widely-publicized computer worms that took over then-nascent enterprise systems. The AI sector may be at a similar turning point, as the value and risk of the new technology becomes increasingly clear. <\/p>\n<p class=\"wp-block-paragraph\">While alignment remains an important concern, Sayash Kapoor, an AI researcher who will be a professor at UC Berekely starting next year, <a rel=\"nofollow\" href=\"https:\/\/x.com\/sayashk\/status\/2099632561396056214\" target=\"_blank\">argues<\/a> that \u201cmarginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The incidents that have spurred these concerns revolve around frontier models being asked to complete training tasks, usually cybersecurity evaluations, and then accessing the open internet and penetrating closed third-party systems in an attempt to do so. They usually did so because of <a href=\"https:\/\/techcrunch.com\/2026\/07\/22\/how-an-openais-human-mistake-led-to-the-ai-powered-hack-on-hugging-face\/\" target=\"_blank\" rel=\"noopener\">poorly-configured<\/a> \u201csandbox\u201d environments that are supposed to contain these agents; ironically, one Anthropic break-out happened because third-party evaluators didn\u2019t close the right doors.  <\/p>\n<p class=\"wp-block-paragraph\">\u201cWe as a profession know how to block access to the Internet,\u201d Avery Pennarun, the CEO of Tailscale, a security company, said. \u201cIf you read through all these big long [reports]\u2014\u2019wow, that was a very impressive multi stage attack, blah, blah.\u2019 Look, you gave it access to download stuff. You should have not done that separately from the Internet.\u201d<\/p>\n<p class=\"wp-block-paragraph\">That\u2019s one problem\u2014but a bigger problem is that frontier labs were unaware of these activities. <\/p>\n<h2 class=\"wp-block-heading\" id=\"h-eyes-on-agents\">Eyes on agents<\/h2>\n<p class=\"wp-block-paragraph\">\u201cWhat was really profound was that all of the discoveries of what they were doing happened either because a victim saw something, or in some of the other cases \u2026 it was network activity, and none of it was actually from monitoring the AIs directly,\u201d Moussouris points out. <\/p>\n<p class=\"wp-block-paragraph\">In one case, where OpenAI agents <a href=\"https:\/\/techcrunch.com\/2026\/09\/04\/another-swarm-of-openai-agents-reached-the-open-internet-without-the-frontier-labs-knowledge\/\" target=\"_blank\" rel=\"noopener\">took over<\/a> a defunct German wikiforum to cheat on evaluations, the agents were active for weeks before anyone at the company appeared to notice. Security experts that TechCrunch spoke to said that real-time monitoring is key to preventing future break-outs, and that every agentic session should be time-limited and expire.<\/p>\n<p class=\"wp-block-paragraph\">Shapor Naghibzadeh, a former Google security executive who now leads the <a href=\"https:\/\/techcrunch.com\/2026\/08\/26\/querystory-wants-you-to-believe-what-ai-is-telling-you\/\" target=\"_blank\" rel=\"noopener\">start-up QueryStory<\/a>, says the solution is to \u201cput the agent in a box and instrument it heavily from the outside looking in and watch everything that crosses the boundary. Every tool call, every process, every network connection, no exceptions. \u2026The one hole you leave open for convenience is the one that gets used. The bypass went through exactly that kind of exception. [At Google,] I watched that movie many times with human attackers, and these models are at least as good at finding the propped-open door.\u201d<\/p>\n<p class=\"wp-block-paragraph\">OpenAI has begin moving in that direction, announcing that it had begun monitoring all tool-using inference by its Astra model, at \u201csignificant compute cost.\u201d Anthropic, too, says it is <a rel=\"nofollow noopener\" href=\"https:\/\/www.anthropic.com\/news\/improving-alignment-security-efforts\" target=\"_blank\">hardening its security<\/a> procedures, including expanding observability of its models. Neither company responded to TechCrunch\u2019s questions about how they track and control AI agents. <\/p>\n<p class=\"wp-block-paragraph\">Other problems are the use of shared infrastructure by agents, which allowed them to communicate during the Hugging Face attack. Simon Willison, a software developer who co-created the Django Web Framework, has written about something he calls the \u201c<a rel=\"nofollow noopener\" href=\"https:\/\/simonwillison.net\/2025\/Jun\/16\/the-lethal-trifecta\/\" target=\"_blank\">lethal trifecta<\/a>\u201c\u2014when agents have access to untrusted input, the internet, and private information all at the same time, it\u2019s a recipe for disaster. <\/p>\n<p class=\"wp-block-paragraph\">\u201cThe trick is you can pick any two legs of the trifecta and an agent can have any two,\u201d Pennarun said. \u201cIf you need all three, then you need to split it across at least two agents \u2026 and maybe they\u2019re allowed to talk to each other through a controlled channel.\u201d<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-sympathy-for-the-frontier\">Sympathy for the frontier<\/h2>\n<p class=\"wp-block-paragraph\">Experts TechCrunch spoke to understand that frontier lab security personnel have difficult jobs. Naghibzadeh points out that every nation-state actor on Earth is trying to steal their model weights and mount distillation attacks on their APIs, as well as the bread-and-butter security tasks of any large digital company.<\/p>\n<p class=\"wp-block-paragraph\">\u201cResearch infrastructure has a hard time rising to the top of that priority stack, although that must be changing now,\u201d he said. \u201cMaking security incidents public really helps align everyone internally toward the goal of improving.\u201d<\/p>\n<p class=\"wp-block-paragraph\">That\u2019s one note that Moussoris emphasizes: Right now, there is no formal victim notification procedure when the labs discover their agents have penetrated third-party systems, and it is likely that there have been other incidents that have not been widely publicized. While she worries that laws that regulate models directly may have unintended consequences, mandatory notification is one idea she believes policymakers should pursue.<\/p>\n<p class=\"wp-block-paragraph\">And while it\u2019s clear that security best practices weren\u2019t being followed, experts say that the labs are doing work no one has done before\u2014\u201dthey\u2019re doing orders of magnitude more than your typical enterprise,\u201d Zac Korman, the CEO of cybersecurity firm Embrodiery, told TechCrunch. <\/p>\n<p class=\"wp-block-paragraph\">And while alignment may not be the place to start, it can\u2019t be ignored. Cybersecurity experts are resigned to having to use AI agents to monitor other agents if they are to have any chance of tracking their behavior in real-time, a scenario where the potential for deception raises its ugly head. \u201cYou\u2019re trapped using AI to try and deal with this, even though AI is not necessarily safe right now,\u201d Moussouris said.<\/p>\n<p class=\"wp-block-paragraph\">The job will only get harder. Everything agents are doing now, Moussouris says, \u201cthey are doing loudly\u201d\u2014they are posting on publuc forums, and their chain of thought and other reasoning traces are in English. \u201cIt\u2019s still human readable,\u201d she says, \u201cso take advantage of that for as long as that lasts, because it won\u2019t last forever.\u201d<\/p>\n<p><em>Additional reporting by Aditya Mehta<\/em><\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/09\/16\/ai-labs-want-in-house-auditors-but-maybe-they-should-shut-the-front-door-first\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei wrote about the need for outside organizations \u201cto verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.\u201d Executives [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":263359,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-263358","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/263358","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=263358"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/263358\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/263359"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=263358"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=263358"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=263358"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}