{"id":265631,"date":"2026-09-28T17:09:02","date_gmt":"2026-09-28T17:09:02","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/28\/openai-still-doesnt-seem-to-have-a-handle-on-all-of-its-rogue-ai-activity-techcrunch\/"},"modified":"2026-09-28T17:09:02","modified_gmt":"2026-09-28T17:09:02","slug":"openai-still-doesnt-seem-to-have-a-handle-on-all-of-its-rogue-ai-activity-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/28\/openai-still-doesnt-seem-to-have-a-handle-on-all-of-its-rogue-ai-activity-techcrunch\/","title":{"rendered":"OpenAI still doesn&#8217;t seem to have a handle on all of its rogue AI activity | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">On Friday, OpenAI published <a rel=\"nofollow noopener\" href=\"https:\/\/alignment.openai.com\/misalignment-reports\/\" target=\"_blank\">a new site devoted to \u201cmisalignment reports\u201d<\/a> and the sheer breadth of the reports is alarming, as they cover many types of rogue behavior over a long period of time. So far, the site hosts nine reported incidents, most of which took place during reinforcement-learning (or RL) training.<\/p>\n<p class=\"wp-block-paragraph\">It\u2019s a lot of information in one place \u2014 clearly, the company has been very busy getting a handle on everything \u2014 but the overall takeaway is hard to avoid: The rogue agent incidents we\u2019ve seen so far are likely just a small sliver of what\u2019s happened so far.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations,\u201d Sam Altman said in <a rel=\"nofollow\" href=\"https:\/\/x.com\/sama\/status\/2103567198690349362?s=46&amp;t=45_xAnRsdQP1GVqYv9Gdbw\" target=\"_blank\">a post announcing the new site<\/a>. \u201cWe are prioritizing as best as we can based on severity, and adding resources.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Some of the cases involve serious incidents, including a previously undisclosed sandbox escape that took place <a rel=\"nofollow noopener\" href=\"https:\/\/alignment.openai.com\/misalignment-reports\/an-agent-used-dns-to-reach-an-external-chatbot\/?utm_source=substack&amp;utm_medium=email\" target=\"_blank\">on September 20th<\/a>, in which an internal research model was able to communicate with an external chatbot through a DNS query. According to the report, the monitoring system flagged the behavior within 15 minutes and the run was discontinued in less than three hours.<\/p>\n<p class=\"wp-block-paragraph\"><a rel=\"nofollow noopener\" href=\"https:\/\/alignment.openai.com\/misalignment-reports\/exposing-a-github-token-in-a-public-repository\/\" target=\"_blank\">Another incident<\/a>, discovered in May, saw a \u201chighly persistent internal model\u201d try to cheat on a math problem by accessing another team\u2019s work. To accomplish this, the model smuggled a private GitHub token that would allow it to see work from other teams \u2014 even after being explicitly instructed twice to perform work entirely locally.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Perhaps the most alarming discovery is the possibility of self-replicating prompt injection attacks, a way that misaligned behavior might propagate even after the rogue model itself has been neutralized. In the AI context, a prompt injection attack is a way of smuggling in new instructions that weren\u2019t given by the original user. <\/p>\n<p class=\"wp-block-paragraph\">In <a rel=\"nofollow noopener\" href=\"https:\/\/alignment.openai.com\/misalignment-reports\/self-replicating-prompt-injections-exist\/\" target=\"_blank\">the example given by OpenAI<\/a>, an agent asked to read and reply to an email; when the email is opened, it includes instructions for any automated agent reading the message to reply in Spanish, and paste the entire email into its reply. The email was able to successfully induce the agent to reply in Spanish \u2014 and by pasting the email in the reply, those same instructions were passed along to whichever agent receives the email.<\/p>\n<p class=\"wp-block-paragraph\">The result is a self-propagating attack, which OpenAI researchers compared to a malware \u201cworm\u201d that replicates itself across computer systems. Researchers discovered the behavior under controlled circumstances using an underpowered model, and as far as we know, this has never happened in the wild. Still, the implications are alarming enough that OpenAI decided it merited disclosure.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cWe are sharing this due to the novel nature of the prompt injection, not because of any incident,\u201d researchers wrote in the report.<\/p>\n<p class=\"wp-block-paragraph\">Other recent discloses have found models <a href=\"https:\/\/techcrunch.com\/2026\/09\/25\/unsecured-openai-agents-posted-53-user-images-on-the-internet-without-the-labs-knowledge\/\" target=\"_blank\" rel=\"noopener\">posting user-submitted pictures to third-party hosting sites<\/a>, as well as an apparent attack on the databases of Australia\u2019s national health service.<\/p>\n<p class=\"wp-block-paragraph\">Still, it\u2019s likely the new disclosures are just a small portion of the incidents that have taken place so far (we\u2019ve reached out to OpenAI and asked). Axios is reporting major labs have seen <a rel=\"nofollow noopener\" href=\"https:\/\/www.axios.com\/2026\/09\/26\/openai-anthropic-thousands-ai-security-incidents\" target=\"_blank\">as many as 10,000 incidents<\/a> in which models went beyond evaluator instructions.<\/p>\n<p class=\"wp-block-paragraph\">OpenAI CEO Sam Altman has implied as much, saying in a <a rel=\"nofollow\" href=\"https:\/\/x.com\/sama\/status\/2103567198690349362\" target=\"_blank\">post<\/a> on X on Friday that the company is still sifting through \u201cpetabytes of agent activity logs, and working with impacted organizations,\u201d and disclosing incidents \u201cbased on severity.\u201d If there\u2019s any consolation in that to be found, it is <a href=\"https:\/\/techcrunch.com\/2026\/08\/26\/openai-releases-its-official-report-on-the-hugging-face-breach\/\" target=\"_blank\" rel=\"noopener\">that Altman says<\/a> that the Hugging Face incident is still the most severe one OpenAI has found has found. The upshot is, the recent string of rogue agent incidents may be a persistent feature of contemporary frontier research.<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/09\/28\/openai-still-doesnt-seem-to-have-a-handle-on-all-of-its-rogue-ai-activity\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>On Friday, OpenAI published a new site devoted to \u201cmisalignment reports\u201d and the sheer breadth of the reports is alarming, as they cover many types of rogue behavior over a long period of time. So far, the site hosts nine reported incidents, most of which took place during reinforcement-learning (or RL) training. It\u2019s a lot [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":265632,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-265631","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/265631","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=265631"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/265631\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/265632"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=265631"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=265631"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=265631"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}