Issue #13
WEDNESDAY 19 AUGUST 2026 · 8 CLAIMS CHECKED · 0 SURVIVED THE RECEIPTS · ISSUE 13 OF 30
- TRUE, BUT · Rented HaloAI is building AI, the headlines say. So independent researchers handed frontier agents real, unpublished research questions and six days each. The agents did all of the engineering and wrote up the results. The papers' own authors rejected both. One got a Strong Reject.
- TRUE, BUT · Cherry-Picked SliceGPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.
- TRUE, BUT · Zero UnderneathOpenAI now predicts your age and your despair. It has published the accuracy of neither.
- TRUE, BUT · Rented HaloClaude hit 14 of 15 protein targets, and outside labs confirmed it. Then read the method: Claude drove the specialist design tools the field already ships, and a binder is the first step of a drug, not the drug.
- TRUE, BUT · Narrowed SuperlativeClaude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.
- TRUE, BUT · Moved RulerTwo new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.
- TRUE, BUT · Lab Not FieldOpenAI and Anthropic are selling the same next step for AI agents: more of them. Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.
- TRUE, BUT · Moved RulerThe coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.
You just read the free check. Members get the full autopsy: evidence trail, steelman, every source. 1 month free.START 30 DAYS FREE ↗