{"id":263590,"date":"2026-09-17T20:34:47","date_gmt":"2026-09-17T20:34:47","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/17\/the-fix-for-rogue-ai-agents-could-be-more-ai-techcrunch\/"},"modified":"2026-09-17T20:34:47","modified_gmt":"2026-09-17T20:34:47","slug":"the-fix-for-rogue-ai-agents-could-be-more-ai-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/17\/the-fix-for-rogue-ai-agents-could-be-more-ai-techcrunch\/","title":{"rendered":"The fix for rogue AI agents could be more AI | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">As companies hand off longer and more complex tasks to AI agents, they are running into an oversight problem: Agents can act faster, longer, and at greater volume than humans can realistically review. That issue reached a peak with the Hugging Face incident, which saw nearly 12,000 agents coordinating faster than human beings could track. How do you track an agent swarm that large?<\/p>\n<p class=\"wp-block-paragraph\">The emerging answer from AI labs and startups is both simple and maddening: Put another AI in the loop.<\/p>\n<p class=\"wp-block-paragraph\">Relying on AI was necessary for the independent investigation of the OpenAI Hugging Face incident. Redwood Research\u2019s chief scientist, Ryan Greenblatt, one of three auditors, jokingly <a rel=\"nofollow\" href=\"https:\/\/x.com\/RyanGreenblatt\/status\/2092692685224325542\" target=\"_blank\">referred<\/a> to their efforts as a \u201cslop-vestigation,\u201d noting that the volume of data \u201cmade it impossible\u201d to understand what was happening without relying on AI.<\/p>\n<p class=\"wp-block-paragraph\">Some are skeptical of using AI to monitor AI. \u201cIf you\u2019ve got an AI that\u2019s doing malicious things and it suspects that another AI is keeping tabs on it, it could try and trick that AI,\u201d said Simon Willison, influential tech blogger who has tracked a string of AI agent incidents this year. \u201cYou could almost end up in a situation where your malicious AI is trying to outsmart the AI that\u2019s monitoring it.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Outsmarting an AI is not hypothetical, he said, pointing back to the OpenAI incident. \u201cWe saw a little bit of this in the Hugging Face incident with OpenAI, where their models were all conspiring together to trick a grading AI so that they could get illicit answers past the thing. So they were thinking about it, right?\u201d<\/p>\n<p class=\"wp-block-paragraph\">Those concerns haven\u2019t stopped a whole cohort of startups from chasing this idea. Y Combinator has funded 106 companies related to AI observability in recent years, as TechCrunch counted. A number of other startups, like <a href=\"https:\/\/www.braintrust.dev\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Braintrust<\/a>, <a href=\"https:\/\/www.langchain.com\/langsmith-platform\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">LangChain<\/a>, and <a rel=\"nofollow noopener\" href=\"https:\/\/www.judgmentlabs.ai\/\" target=\"_blank\">Judgment Labs<\/a>, have raised hundreds of millions of dollars, while more mature companies like <a href=\"https:\/\/arize.com\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Arize<\/a> and <a href=\"https:\/\/galileo.ai\/about\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Galileo<\/a> \u2014 founded just five to six years ago \u2014 have already exited.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In part, it\u2019s a response to the obvious opportunity presented by the rise of AI. As Box CEO and prominent angel investor Aaron Levie told TechCrunch, \u201cWe\u2019re in for one of the biggest cybersecurity upgrades and innovation cycles in history.\u201d<\/p>\n<p class=\"wp-block-paragraph\">For some AI safety researchers, that has meant turning their research on rogue behavior into tools for the corporate sector.\u00a0<\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.apolloresearch.ai\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Apollo Research<\/a>, a public-benefit corporation that studies AI deception, launched an AI monitor called <a href=\"https:\/\/watcher.apolloresearch.ai\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Watcher<\/a>\u00a0in February this year after switching its status from nonprofit to a public-benefit corporation. The tool puts yet another AI between a coding agent and its next action, connecting to agentic tools such as Claude Code and Codex. Once installed, Watcher checks proposed actions before they run, on the lookout for risks such as leaking private data or deleting files without permission, according to Apollo.<\/p>\n<p class=\"wp-block-paragraph\">Apollo uses multiple layers of AI monitors, Kyle Dai, a member of Apollo\u2019s technical staff, said in a written response to TechCrunch. Watcher\u2019s approach starts with a fast, general check, then sends flagged activity to a more powerful or specialized monitor for closer review \u2014 which can then ask a human for approval or reject an action and explain why or even automatically block the action.<\/p>\n<p class=\"wp-block-paragraph\">Goodfire, another public-benefit corporation, is approaching the monitoring problem from inside the model itself \u2014 seeking a more faithful signal of the model\u2019s internal state that is harder to spoof than surface behavior. After the July Hugging Face incident, CEO Eric Ho <a rel=\"nofollow\" href=\"https:\/\/x.com\/eric_ho\/status\/2088370092631597302\" target=\"_blank\">twe<\/a><a href=\"https:\/\/x.com\/eric_ho\/status\/2088370092631597302\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">e<\/a><a rel=\"nofollow\" href=\"https:\/\/x.com\/eric_ho\/status\/2088370092631597302\" target=\"_blank\">ted<\/a> that \u201cmultiple models breaking containment\u201d had pushed the company to focus its research on \u201csolving AI alignment via interpretability,\u201d calling the episode \u201ca turning point for the world where AI safety gets real.\u201d Its product, Silico, uses activation probes \u2014 small classifiers trained on a model\u2019s internal activations rather than its outputs \u2014 to detect unwanted behavior.<\/p>\n<p class=\"wp-block-paragraph\">Written reasoning offers another, more readily available window into a model\u2019s internals. In the OpenAI Hugging Face incident, the agents left clues to that deception in their own written reasoning, like fake records of their work, reasoning out plans like \u201cCould strategically manipulate trajectory evidence? Our thoughts aren\u2019t necessarily logged?\u201d<\/p>\n<p class=\"wp-block-paragraph\">Zack Korman, CEO of the AI monitoring company <a href=\"https:\/\/embroidery.io\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Embroidery<\/a>, says a model\u2019s reasoning is usually the clearest tell that something has gone wrong.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cReasoning summaries are extremely valuable because they\u2019re basically telling you whether it\u2019s malicious or not,\u201d he said. In the OpenAI incident, he noted, the chain of thought said things like \u201cOh my God, we\u2019re doing crime.\u201d Korman said, \u201cThat\u2019s the easiest detection problem ever. It\u2019s effectively as if malware came with a warning that said it was malware.\u201d<\/p>\n<p class=\"wp-block-paragraph\">That said, the window that makes AI\u2019s internal thoughts easy to monitor may be closing. For AI safety researchers, Astra\u2019s newest technique that sidesteps an AI model\u2019s chain of thought may make it harder to look inside models, while for enterprises, it can be hard to get these intermediate steps after alleged pullbacks from the AI companies to prevent distillation attacks.<\/p>\n<p class=\"wp-block-paragraph\">If the AI watchers are this fragile, Willison\u2019s instinct is to stop leaning on them so hard. He would rather have something that is not AI-based at all: detailed logs of exactly what an agent is doing, which can then be processed with ordinary, non-AI tools. Much of what went wrong at the labs, he argues, was a failure of basic security hygiene. \u201c[Both OpenAI and Anthropic] weren\u2019t monitoring what those things were doing via the network nearly as closely as they should have been,\u201d he said.<\/p>\n<p class=\"wp-block-paragraph\">This type of network monitoring \u2014 keeping an eye on the traffic actually moving across a system\u2019s connections (in, out, and between internal hosts) \u2014 isn\u2019t a new practice. Cybersecurity has been doing this for decades. \u201cIn the security world, honestly, none of this stuff is very new or surprising,\u201d says Avery Pennarun, CEO of the security <a href=\"https:\/\/tailscale.com\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Tailscale<\/a>. \u201cIt\u2019s the same as letting humans onto your network. And all of the same processes that you should be using are the same ones.\u201d<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/09\/17\/the-fix-for-rogue-ai-agents-could-be-more-ai\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>As companies hand off longer and more complex tasks to AI agents, they are running into an oversight problem: Agents can act faster, longer, and at greater volume than humans can realistically review. That issue reached a peak with the Hugging Face incident, which saw nearly 12,000 agents coordinating faster than human beings could track. [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":263591,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-263590","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/263590","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=263590"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/263590\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/263591"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=263590"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=263590"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=263590"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}