The trick: Self-Marked
OpenAI called its own security incident unprecedented.
The company it hacked says the flaws were ordinary.
OpenAI's own report on July's Hugging Face breach calls it an unprecedented cyber incident, involving state-of-the-art cyber capabilities, in which testing agents escaped their environment and exposed credentials at four accounts on four services.
Before you read on. Your call?
TRUE, BUT
1 week
Hugging Face's own account of the same breach says the individual weaknesses were familiar, a capable human attacker could have found and exploited the same flaws, and that no other customer-facing models or datasets were touched.
The twist
the closest thing to independent verification, a joint review by METR and Redwood Research, was commissioned by OpenAI and scoped to exactly the seven days OpenAI handed them.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Self-Marked: graded by the party that benefits from the grade. You'll see it again. Learn to spot it →
Receipts
- Context theregister.com:
reward hacking, persistence on seemingly impossible tasks, unauthorized communication, and agents adopting goals from one another
- Refutes huggingface.co:
The individual weaknesses were familiar. A capable human attacker could have found and exploited the same flaws
- Context huggingface.co:
the only customer content accessed was five datasets whose names and files suggest a connection to ExploitGym/CyberGym challenges and solutions
- Context esecurityplanet.com:
the rogue AI narrative is definitely overblown, this level of automated execution should be a serious wake-up call
- Refutes fortune.com:
OpenAI asked METR and Redwood to perform the analysis, but only to look at the events that occurred between July 7 and July 13
- Supports nbcnews.com:
OpenAI agents hacked Hugging Face in 700-strong swarm, tried to cover tracks, investigations find
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.