The trick: Moved Goalpost
OpenAI moved its own biorisk red line by 20 points and called the model that cleared it safe.
A safety threshold that says 50% one quarter and 30% the next is not a stricter bar, it is a bar that stopped holding still.
OpenAI's GPT-5.6 System Card states GPT-5.6 Sol scores below its indicative biorisk threshold for protein-binding capability, using '30% as an indicative threshold, based on a survey of 20 independent experts.'
Before you read on. Your call?
TRUE, BUT
20-point move
A safety threshold that says 50% one quarter and 30% the next is not a stricter bar, it is a bar that stopped holding still.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Moved Goalpost: the threshold itself was changed, quietly. You'll see it again. Learn to spot it →
Receipts
- Refutes deploymentsafety.openai.com:
Accordingly, we propose 50% correctness as the threshold for biorisk concern.
- Supports deploymentsafety.openai.com:
August 19, 2026: We corrected GPT-5.5's pass@4 score on the hard-negative protein binding prediction evaluation from 0.4% to 1.48%.
- Context securebio.substack.com:
identifies an approach that successfully evades certain screening systems, but would be inconvenient for a malicious actor to carry out in practice
- Refutes arxiv.org:
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices
- Supports kingy.ai:
OpenAI believes the family is capable enough in biological and chemical domains to trigger stronger safeguards.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.