{"id":267595,"date":"2026-10-08T16:00:00","date_gmt":"2026-10-08T16:00:00","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/10\/08\/goodfire-says-its-new-inside-out-monitors-catch-rogue-ai-agents-at-a-fraction-of-the-cost-techcrunch\/"},"modified":"2026-10-08T16:00:00","modified_gmt":"2026-10-08T16:00:00","slug":"goodfire-says-its-new-inside-out-monitors-catch-rogue-ai-agents-at-a-fraction-of-the-cost-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/10\/08\/goodfire-says-its-new-inside-out-monitors-catch-rogue-ai-agents-at-a-fraction-of-the-cost-techcrunch\/","title":{"rendered":"Goodfire says its new \u2018inside-out\u2019 monitors catch rogue AI agents at a fraction of the cost | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">The standard way to keep an AI agent in line is to <a href=\"https:\/\/techcrunch.com\/2026\/09\/17\/the-fix-for-rogue-ai-agents-could-be-more-ai\/\" target=\"_blank\" rel=\"noopener\">have a second AI read over its shoulder<\/a>. It\u2019s been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels\u2019 worth of text.<\/p>\n<p class=\"wp-block-paragraph\">Goodfire, a startup focused on interpretability (figuring out how AI models work internally), launched a cheaper option on Thursday: monitors that watch what\u2019s happening inside an AI model as it works, rather than just reading what it writes. The monitors are available to customers of Baseten, which hosts and runs AI models for other companies.<\/p>\n<p class=\"wp-block-paragraph\">Baseten\u2019s Base Labs <a href=\"https:\/\/techcrunch.com\/2026\/09\/17\/base-labs-launches-an-open-weight-ai-safety-partnership-with-hugging-face-and-goodfire\/\" target=\"_blank\" rel=\"noopener\">announced a safety partnership<\/a> with Goodfire and the AI platform Hugging Face last month.<\/p>\n<p class=\"wp-block-paragraph\">The launch comes after <a href=\"https:\/\/techcrunch.com\/2026\/08\/27\/heres-all-the-times-ai-has-gone-rogue-and-hacked-other-companies\/\" target=\"_blank\" rel=\"noopener\">a string of incidents this year<\/a> in which AI agents escaped their test environments, including OpenAI agents that <a href=\"https:\/\/techcrunch.com\/2026\/07\/27\/openais-hugging-face-breach-has-reignited-the-debate-over-alignment-and-control\/\" target=\"_blank\" rel=\"noopener\">breached Hugging Face<\/a>. Kimi K3, the open model Goodfire built its first monitor around, <a href=\"https:\/\/techcrunch.com\/2026\/08\/07\/chinese-ai-model-kimi-escaped-its-cybersecurity-testing-environment-researchers-say\/\" target=\"_blank\" rel=\"noopener\">took advantage of a leak in its sandbox<\/a> to access the internet and information on GitHub this summer.<\/p>\n<p class=\"wp-block-paragraph\">Goodfire\u2019s system works a bit like airport security. Small detectors called probes read the model\u2019s internal signals at every step of an agent\u2019s work, the way a walk-through scanner checks every passenger. Only when a probe flags something does a separate AI model, the equivalent of a hand search, take a closer look.<\/p>\n<p class=\"wp-block-paragraph\">Baseten customers can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking. They also decide the automated response: logging the event, sending it for human review, or refusing the request entirely.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Goodfire says its approach is also cheaper to run. Most AI monitors are separate models that have to reread everything the monitored model does, which adds time and cost. Goodfire\u2019s probes instead tap into calculations the model is already making as it works.<\/p>\n<p class=\"wp-block-paragraph\">\u201cInternal activation monitors are really cheap because they reuse the computations in the forward pass,\u201d Goodfire CEO Eric Ho <a rel=\"nofollow noopener\" href=\"https:\/\/www.mattturck.com\/podcast\/why-ai-agents-cheat-eric-ho-goodfire\" target=\"_blank\">said<\/a> on venture capitalist Matt Turck\u2019s MAD Podcast last week. \u201cSo the model\u2019s already computing this token. All you\u2019re doing is taking the intermediate neural activations that it\u2019s already computed and then running a classifier over these internal computations.\u201d In short, the model is already doing the math, and the probes just read the results.<\/p>\n<p class=\"wp-block-paragraph\">In Goodfire\u2019s tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51, compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier one. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look.<\/p>\n<p class=\"wp-block-paragraph\">Running four probes at once added less than 2% to the time it takes the model to start responding, the company said.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe great advantage is that you can catch things before they happen,\u201d Goodfire CTO and co-founder Dan Balsam said. \u201cWe can detect when the model <em>might<\/em> hack during eval or training.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The pitch is aimed at open models. Developers can download them and <a href=\"https:\/\/techcrunch.com\/2026\/09\/03\/abliteration-ai-is-making-a-business-out-of-removing-ai-guardrails\/\" target=\"_blank\" rel=\"noopener\">strip out their safeguards<\/a>, and they don\u2019t come with the kind of <a href=\"https:\/\/techcrunch.com\/2026\/08\/26\/openai-releases-its-official-report-on-the-hugging-face-breach\/\" target=\"_blank\" rel=\"noopener\">monitoring that closed labs run<\/a> on their own systems.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe damage that an individual can do with an open model is small compared to what someone can do with clusters of compute, like inference providers\u2014where most of the liability is,\u201d said Balsam. \u201cWhen we have the open \u201cMythos\u201d moment, it\u2019s going to become clear that models need guardrails deployed at inference time.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Goodfire\u2019s <a rel=\"nofollow noopener\" href=\"https:\/\/www.goodfire.com\/research\/reward-hacking-activation-monitors#\" target=\"_blank\">recent research<\/a> found that leading open models, including Kimi K3 and GLM 5.2, reward-hacked in 50% to 96% of runs on tests of AI agents.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Goodfire isn\u2019t the first to try this approach. Google DeepMind said in January that its research informed the deployment of <a rel=\"nofollow noopener\" href=\"https:\/\/arxiv.org\/abs\/2601.11516\" target=\"_blank\">misuse-detection probes in Gemini<\/a>.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Balsam said the monitors are the near-term piece of a longer research goal: reverse-engineering an LLM so that behavior can be traced back to where it emerged in training. \u201cWe hope to turn the magic of training models into precision engineering, \u201d he said.<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/10\/08\/goodfire-says-its-new-inside-out-monitors-catch-rogue-ai-agents-at-a-fraction-of-the-cost\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The standard way to keep an AI agent in line is to have a second AI read over its shoulder. It\u2019s been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels\u2019 worth of text. Goodfire, a startup focused on interpretability (figuring out how [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":267596,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-267595","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/267595","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=267595"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/267595\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/267596"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=267595"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=267595"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=267595"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}