{"id":253354,"date":"2026-07-24T01:00:00","date_gmt":"2026-07-24T01:00:00","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/07\/24\/how-ai-guardrails-are-impeding-the-work-of-offensive-cybersecurity-researchers-techcrunch\/"},"modified":"2026-07-24T01:00:00","modified_gmt":"2026-07-24T01:00:00","slug":"how-ai-guardrails-are-impeding-the-work-of-offensive-cybersecurity-researchers-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/07\/24\/how-ai-guardrails-are-impeding-the-work-of-offensive-cybersecurity-researchers-techcrunch\/","title":{"rendered":"How AI guardrails are impeding the work of offensive cybersecurity researchers | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">For months, AI giants have devised special vetted programs and strict guardrails to limit the use of their models by malicious hackers. But these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In June, the U.S. government <a href=\"https:\/\/techcrunch.com\/2026\/06\/12\/anthropics-safety-warnings-may-have-just-backfired-the-government-has-pulled-the-plug-on-its-most-powerful-ai\/\" target=\"_blank\" rel=\"noopener\">slapped export control restrictions<\/a> on Anthropic\u2019s much-hyped AI models Mythos and Fable. The move was prompted at least in part by a report that claimed it was possible to bypass the models\u2019 guardrails designed to prevent users from using them to build and execute malicious cyberattacks.<\/p>\n<p class=\"wp-block-paragraph\">Regardless of whether the incident was really motivated by <a href=\"https:\/\/techcrunch.com\/2026\/06\/15\/the-us-governments-anthropic-models-ban-was-never-about-an-ai-jailbreak\/\" target=\"_blank\" rel=\"noopener\">fears of a jailbreak<\/a>, the fact is that Anthropic <a href=\"https:\/\/techcrunch.com\/2026\/04\/07\/anthropic-mythos-ai-model-preview-security\/\" target=\"_blank\" rel=\"noopener\">has repeatedly marketed<\/a> Mythos as <a href=\"https:\/\/techcrunch.com\/2026\/04\/09\/is-anthropic-limiting-the-release-of-mythos-to-protect-the-internet-or-anthropic\/\" target=\"_blank\" rel=\"noopener\">some kind of doomsday cybermachine<\/a> that can only be given to carefully vetted users, and even then with strict guardrails in place. (The export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1; Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government\u2019s review process.)<\/p>\n<p class=\"wp-block-paragraph\">That kind of gatekeeping isn\u2019t unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs they can apply to get vetted and \u2014 if approved \u2014 access models with fewer cybersecurity restrictions: OpenAI\u2019s <a rel=\"nofollow noopener\" href=\"https:\/\/openai.com\/index\/scaling-trusted-access-for-cyber-defense\/\" target=\"_blank\">Trusted Access for Cyber<\/a> and Anthropic\u2019s <a rel=\"nofollow noopener\" href=\"https:\/\/support.claude.com\/en\/articles\/14604842-real-time-cyber-safeguards-on-claude-opus-and-sonnet\" target=\"_blank\">Cyber Verification Program<\/a>.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">These guardrails have been widely criticized, particularly by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.<\/p>\n<p class=\"wp-block-paragraph\">During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, <a rel=\"nofollow noopener\" href=\"https:\/\/docs.google.com\/document\/d\/1G2B7VetSNfxN9Lfb8f2Y1Vyy-uVBQPl_YSj_3ayVUxM\/edit?tab=t.0\" target=\"_blank\">said<\/a> that, \u201cit\u2019s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what\u2019s not.\u201d <\/p>\n<p class=\"wp-block-paragraph\">Dowd has spent decades <a rel=\"nofollow noopener\" href=\"https:\/\/www.vice.com\/en\/article\/iphone-zero-days-inside-azimuth-security\/\" target=\"_blank\">finding and selling \u201czero days\u201d<\/a> \u2014 previously unknown software flaws and the exploits that take advantage of them \u2014 to Western governments, rather than report them to the software makers so they get patched. Governments pay a premium for vulnerabilities precisely because they stay open, which is useful for intelligence operations.<\/p>\n<p class=\"wp-block-paragraph\">Dowd admitted his work may make him biased, but he isn\u2019t alone. Several people who work in offensive cybersecurity \u2014 they proactively probe systems for weaknesses  \u2014 described to TechCrunch how they use AI tools and deal with their guardrails.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Chris Anley, the chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it\u2019s a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders, he said. <\/p>\n<p class=\"wp-block-paragraph\">\u201cThis is where the whole offensive versus defensive and guardrails part comes in, because \u2018fix this code\u2019 as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,\u201d said Anley. \u201cSo at the same time, the same tool is both an offensive tool and a defensive tool, and the two can\u2019t really be unpicked.\u201d <\/p>\n<p class=\"wp-block-paragraph\">It\u2019s \u201clike a hammer,\u201d he continued. \u201cYou can\u2019t build a house without a hammer. It\u2019s definitely a tool but it\u2019s also irreducibly a weapon as well.\u201d<\/p>\n<p class=\"wp-block-paragraph\">When he and his colleagues run into such a roadblock, they sometimes fall back on open-source AI models that come with no guardrails at all. <\/p>\n<p class=\"wp-block-paragraph\">Paolo Stagno, the chief technology officer at CrowdFense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying AI companies \u201cessentially treat customers like children who need babysitting\u201d with their vetted programs and guardrails.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Stagno said he and his colleagues do use frontier models \u2014 but only for reverse engineering. They avoid using AI to help find vulnerabilities or build exploits, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, he said, they use open source models run locally, as they do not rely on sharing data outside of the model.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails are not impeding his work. That\u2019s because he doesn\u2019t use AI for offensive work; instead, he uses it for initial reverse engineering, to understand the code he\u2019s analyzing, and to build supporting tools. For that, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cI still want to own the actual bug discovery and weaponization myself and that wouldn\u2019t change if all guardrails were lifted tomorrow,\u201d said Cali. \u201cI am jealous of my bugs, and I like this game too much to let models play it for me.\u201d<\/p>\n<p class=\"wp-block-paragraph\">One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity because he isn\u2019t authorized to talk to the press, said his employer isn\u2019t part of Anthropic\u2019s CVP program and as a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIf it catches wind we\u2019re doing anything security related, it just stops and isn\u2019t usable,\u201d the person said.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Chris Thompson \u2014 chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive security and AI-focused event \u2014 said that in his experience using the frontier AI models, the guardrails can be inconsistent and work differently every day. That\u2019s true even inside the looser boundaries of Anthropic and OpenAI\u2019s vetted programs.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cI think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,\u201d said Thompson. \u201cInstead of analyzing a vulnerability and reasoning through the exploitability, you\u2019re trying to find why you\u2019re getting inconsistent results or why are models over-sanitizing the output.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">As a consequence, researchers rely on or get pushed toward Chinese open-source models like GLM \u2014 freely downloadable models that can be run locally with no vetting or usage restrictions \u2014 said Thompson.<\/p>\n<p class=\"wp-block-paragraph\">\u201cYou have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,\u201d he said. \u201cI think it\u2019s more harmful than good to have these guardrails in place.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Rather than tightening restrictions further, Thompson called for the AI frontier labs to open up their programs, provide responsible access, and also hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThere\u2019s this big storm coming. There\u2019s this big wave of attacks that are going to happen at speed and scale like never before,\u201d said Thompson. \u201cBut the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.\u201d<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/07\/23\/how-ai-guardrails-are-impeding-the-work-of-offensive-cybersecurity-researchers\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>For months, AI giants have devised special vetted programs and strict guardrails to limit the use of their models by malicious hackers. But these limits are now hindering the work of legitimate network defenders, as well as that of offensive cybersecurity researchers.\u00a0 In June, the U.S. government slapped export control restrictions on Anthropic\u2019s much-hyped AI [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":253355,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-253354","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/253354","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=253354"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/253354\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/253355"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=253354"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=253354"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=253354"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}