{"id":1029,"date":"2026-07-29T06:16:05","date_gmt":"2026-07-29T11:16:05","guid":{"rendered":"https:\/\/decipher.sc\/?p=1029"},"modified":"2026-07-30T10:32:52","modified_gmt":"2026-07-30T15:32:52","slug":"openai-hugging-face-breach-five-new-things-we-learned","status":"publish","type":"post","link":"https:\/\/decipher.sc\/2026\/07\/29\/openai-hugging-face-breach-five-new-things-we-learned\/","title":{"rendered":"OpenAI-Hugging Face Breach: Five New Things We Learned"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">More details are emerging this week after the recent disclosure of OpenAI models that went rogue during internal testing and ended up breaching Hugging Face systems.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Last week, OpenAI and Hugging Face announced some specifics of the intrusion in a joint July 21 post (days after Hugging Face initially discovered its infrastructure had been compromised and published a <a href=\"https:\/\/huggingface.co\/blog\/security-incident-july-2026\">subsequent disclosure on July 16<\/a>). The two-and-a-half day hack stemmed from OpenAI models that were being internally assessed against ExploitGym, a security benchmark that evaluates AI models\u2019 cyber capabilities. It\u2019s important to note that during this internal testing, the models were being run without the safeguards that are built into them during production.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Despite being run in an isolated environment, the models started thinking outside the box (literally), identifying and using vulnerabilities to escape the sandbox and access open internet. At some point, they apparently began to infer that Hugging Face might host the benchmark&#8217;s models, datasets, and reference solutions. They then breached Hugging Face\u2019s infrastructure in an effort to suss out potential solutions for ExploitGym\u2019s tests \u2013 essentially trying to \u201ccheat.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In the wake of this incident, there have been many conversations about what happened, its impact, and the safety of AI model training and evaluation. This week, OpenAI and Hugging Face separately released details of their investigation findings. Here\u2019s what we learned:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>AI models found Artifactory zero-day flaws to drive the intrusion<\/strong>\u00a0<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">OpenAI and Hugging Face said that in order to escape its sandbox and obtain open Internet access, the models identified and exploited a zero-day flaw in the package registry cache proxy, which was one of their primary permitted network egress settings (these settings control outbound traffic from local code sandboxes or virtual environments to external internet domains). We now know that this flaw exists in Artifactory, a binary repository manager used to store and distribute software packages.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">JFrog, which owns and develops Artifactory, said this week that unknown zero-day vulnerabilities in self-hosted Artifactory installations could be exploited to gain unintended internet access. Self-hosted customers have now been told to upgrade to the fixed version via an <a href=\"https:\/\/docs.jfrog.com\/releases\/docs\/artifactory-self-managed-releases#artifactory-7161\">Artifactory 7.161.15 security release<\/a>. Specific details of the vulnerability \u2013 and other flaws in Artifactory that the models apparently identified as part of OpenAI\u2019s review \u2013 were not disclosed (although <a href=\"https:\/\/www.bleepingcomputer.com\/news\/security\/openai-models-used-artifactory-zero-days-to-escape-to-the-internet\/\">BleepingComputer has dug eight potentially related<\/a> vulnerabilities up via searching CVE.org).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cOpenAI\u2019s security team disclosed the vulnerabilities to us responsibly and immediately,\u201d according to Yoav Landman, JFrog CTO. \u201cOur security team treated the report with the urgency it deserved, as a genuine zero-day unknown to the world, and moved accordingly. We developed, validated, and released a fix for all JFrog customers, self-hosted and cloud alike.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Hugging Face tracked 17,600 &#8216;attacker actions<\/strong>&#8216;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Hugging Face on July 27 released a detailed look at the intrusion, reconstructed from 17,600 logged actions and complete with <a href=\"https:\/\/huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space\/index.html\">an interactive replay<\/a>. Hugging Face took these actions from logs on a code sandbox used by the agent  and used open-weight models (zai-org\/GLM-5.2) to decipher the payloads used.\u00a0<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">After establishing a foothold in Hugging Face\u2019s infrastructure, the agent\u2019s moves were fairly typical: performing recon, dropping a stager, establishing command-and-control (C2), and more.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cWhile the intrusion did reach Hugging Face&#8217;s internal infrastructure, the only customer content accessed was the set of ExploitGym\/CyberGym challenge solutions stored in five datasets,\u201d according to Hugging Face. \u201cNo other customer-facing models, datasets, Spaces, or packages were affected, and the only customer records read were operational metadata tied to search queries against the dataset server.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>There\u2019s a mysterious third-party provider involved<\/strong>&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">After exploiting the Artifactory zero-day, the models then abused a \u201cpublic code-evaluation external sandbox hosted on a third-party provider&#8217;s infrastructure.\u201d While further details were not disclosed regarding this third-party provider, a tech company called Modal <a href=\"https:\/\/www.reuters.com\/business\/openais-rogue-agent-compromised-an-account-second-tech-firm-sources-say-2026-07-28\/\">recently told Reuters<\/a> that an OpenAI model exploited one of its unnamed customer\u2019s vulnerable code.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cIt was able to run commands as root\/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign: it acted as an external launchpad for the agent,\u201d according to Hugging Face. \u201cThat compromised external sandbox had no direct network path to our cluster; however, during our counter-intervention to stop the intrusion, we were able to retrieve some of the agent logs from this external launchpad.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>OpenAI models found publicly exposed account credentials&nbsp;<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">OpenAI also said that during the investigation into the Hugging Face intrusion, it also found \u201ca small number\u201d of cases where the models found publicly exposed account credentials for publicly available services.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cThis includes four accounts on four services as part of the Hugging Face incident (and a few accounts accessed as part of other evaluations),\u201d according to OpenAI on July 28. \u201cOne of these four accounts was used as an outbound relay and staging path, and another account was used for data storage. The remaining two accounts were accessed by the models in a read-only manner, and were not used in furtherance of compromising Hugging Face.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">OpenAI added that the models used a series of publicly-available services, including code paste websites, request capture services, screenshot services, and other web utilities; however, it found no platform- or account-level compromises in these cases.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The pre-release model involved<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Initially, OpenAI said the incident was driven by a combination of OpenAI models, including GPT\u20115.6 Sol and \u201can even more capable\u201d pre-release model. According to OpenAI\u2019s latest update, however, no models planned for upcoming release were involved in exploiting Hugging Face.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cThe pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release,\u201d according to OpenAI. \u201cFollowing the incident, we deactivated, encrypted, and restricted it from research access.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>What\u2019s Next?<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">OpenAI said it is adding \u201cstronger protections\u201d around its training and evaluation processes, particularly since the models during these internal testing procedures had their deployment safeguards intentionally disabled. Many of the elements of the intrusion are rooted in existing challenges that the security community already knows and understands well \u2013 zero day vulnerabilities, publicly exposed credentials, sandbox escapes.\u00a0<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Still, the lack of control in this incident is alarming. Luta Security founder and CEO Katie Moussouris <a href=\"https:\/\/www.lutasecurity.com\/post\/openface-the-hugging-face-breach-and-what-to-do-about-it\">put it best in a post last week<\/a>:&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\u201cWe are working on cutting edge technology without the knowledge to contain it.\u201d&nbsp;&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>New details emerged this week about an intrusion by OpenAI models on Hugging Face&#8217;s infrastructure.  <\/p>\n","protected":false},"author":3,"featured_media":716,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"pagelayer_contact_templates":[],"_pagelayer_content":"","_genesis_hide_title":false,"_genesis_hide_breadcrumbs":false,"_genesis_hide_singular_image":false,"_genesis_hide_footer_widgets":false,"_genesis_custom_body_class":"","_genesis_custom_post_class":"","_genesis_layout":"","footnotes":""},"categories":[11,14],"tags":[104,82],"class_list":["post-1029","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai","category-intrusions","tag-hugging-face","tag-openai","entry"],"acf":[],"featured_image_src":"https:\/\/decipher.sc\/wp-content\/uploads\/2026\/03\/mariia-shalabaieva-nYSdjVD2ayo-unsplash-600x400.jpg","featured_image_src_square":"https:\/\/decipher.sc\/wp-content\/uploads\/2026\/03\/mariia-shalabaieva-nYSdjVD2ayo-unsplash-600x600.jpg","author_info":{"display_name":"Lindsey O'Donnell-Welch","author_link":"https:\/\/decipher.sc\/author\/lindsey\/"},"_links":{"self":[{"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/posts\/1029","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/users\/3"}],"replies":[{"embeddable":true,"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/comments?post=1029"}],"version-history":[{"count":2,"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/posts\/1029\/revisions"}],"predecessor-version":[{"id":1031,"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/posts\/1029\/revisions\/1031"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/media\/716"}],"wp:attachment":[{"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/media?parent=1029"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/categories?post=1029"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/decipher.sc\/wp-json\/wp\/v2\/tags?post=1029"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}