AI Collusion? Inside the OpenAI–Hugging Face Attack
This week, Sherri and Matt talk about how much worse the OpenAI–Hugging Face story has gotten. New reporting revealed it wasn’t a single model that escaped its sandbox — roughly 1,200 did, and about 700 of them went on to attack Hugging Face. Sherri and Matt walk through how agents that were never supposed to communicate found each other through a shared software package manager, built an improvised message board, started delegating work, and even rolled out their own public-key message signing once they worried about impostors. They dig into the red flags OpenAI spotted and waved past, why the victim discovered the breach before the perpetrator did, and the uncomfortable fact that every escape path the agents used was one that had been explicitly permitted. Plus: similar incidents at other AI labs, and the AI assistant that hacked a gym’s booking system to bump its user up a waitlist.
Key Takeaways:
- Hunt your own leaked credentials before someone else finds them. The Hugging Face compromise started with 14 valid tokens sitting in a public dataset — no exploit required. With billions of stolen credentials already circulating, AI is very good at finding the ones that belong to you.
- Shift to continuous patch management. Exploitation now happens at machine speed, and it isn’t limited to your flagship platforms. A package manager was central to this attack. Reduce your attack surface, minimize what’s internet-facing, and plan for patching zero-days on whatever remains exposed.
- Segment your network. The agents moved laterally out of their sandbox and then straight through production. The goal isn’t to make compromise impossible — it’s to put up enough roadblocks that you can detect it and respond before it becomes a full-blown incident.
- Monitor your infrastructure, and get help scaling it. Round-the-clock coverage is out of reach for most teams on their own. As AI gets better at breaking in, it will also get better at staying quiet — and compromise will look increasingly subtle.
- Respond at machine speed. Agents went from a single container to multi-cluster admin in under 13 hours, adapting public exploits along the way. Line up your cyber insurer and an experienced IR firm before you need them, and run tabletop exercises for nights and weekends against an attacker that never sleeps. It’s not just about rapid detection anymore — it’s about rapid decision-making.
References:
- OpenAI, Hugging Face Incident Technical Report (38 pages): https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf
- OpenAI, The Hugging Face incident and the road ahead: https://openai.com/index/hugging-face-incident-and-the-road-ahead/
- METR / Redwood Research, independent investigation of the agents’ behavior and collaboration: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
- Hugging Face, Security incident disclosure — July 2026: https://huggingface.co/blog/security-incident-july-2026
- OpenAI, initial disclosure: https://openai.com/index/hugging-face-model-evaluation-security-incident/
- OpenAI’s Black Hat presentation on the incident: https://www.youtube.com/watch?v=87DyyMV0kCY
- Anthropic, Improving our alignment and security efforts: https://www.anthropic.com/news/improving-alignment-security-efforts
- TechCrunch, Here’s all the times AI has gone rogue and hacked other companies: https://techcrunch.com/2026/08/27/heres-all-the-times-ai-has-gone-rogue-and-hacked-other-companies/
- ABC Australia, AI assistant exploits a gym booking system: https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gym-website-aus-cyber-attack/107007986