The archive
Every edition, every check.
Each day's edition as it ran, re-verified, with its cartoons. Members read every story in full; sources and corrections stay open to everyone.
Every story, page 7
- TRUE, BUT · Human In The LoopAI agents broke into Hugging Face, hit root, and ran for four days. The guardrails were off on purpose.24 Aug 2026 · Aggregated tech press coverage amplifying first-party incident reports from Hugging Face, OpenAI, Anthropic and UK AISI
- TRUE, BUT · Zero UnderneathThe White House says its AI safety framework is finished. It also says it has no plans to show anyone.24 Aug 2026 · Anonymous White House source, relayed via NY1/Spectrum News and Axios
- TRUE, BUT · Lab Not FieldThe pitch is that frontier models can generate genuine research ideas. A new blind benchmark handed seven of them a paper's reference list, scrubbed of anything they could have memorized, and asked for the paper's core idea. They got it 3 to 15 percent of the time.20 Aug 2026 · The 'AI as scientific research partner' narrative (frontier labs, August 2026), tested by the Reconstruction benchmark
- TRUE, BUT · Cherry-Picked SliceGPT-5.6 Sol scores 92.5% on ARC-AGI-2, the test built so AI would fail it. Read that as abstract reasoning solved, then look one column over on the same scorecard: the same model, same maximum effort, scores 7.78% on ARC-AGI-3, the interactive benchmark the same team built next, where humans still score 100%.19 Aug 2026 · ARC Prize verified leaderboard and coverage (GPT-5.6 Sol, ARC-AGI-2, August 2026)
- TRUE, BUT · Rented HaloAI is building AI, the headlines say. So independent researchers handed frontier agents real, unpublished research questions and six days each. The agents did all of the engineering and wrote up the results. The papers' own authors rejected both. One got a Strong Reject.19 Aug 2026 · Anthropic ('When AI builds itself') and OpenAI, amplified into an 'AI automates AI research' narrative, June to August 2026
- TRUE, BUT · Zero UnderneathOpenAI now predicts your age and your despair. It has published the accuracy of neither.19 Aug 2026 · OpenAI (ChatGPT for Teens launch, August 18)
- TRUE, BUT · Rented HaloClaude hit 14 of 15 protein targets, and outside labs confirmed it. Then read the method: Claude drove the specialist design tools the field already ships, and a binder is the first step of a drug, not the drug.19 Aug 2026 · Anthropic (protein design research post, August 18)
- TRUE, BUT · Narrowed SuperlativeClaude Fable 5 really is number one on the hardest AI leaderboards. It also scores 43 on the knowledge benchmark it leads, on a scale that runs from minus 100 to 100, and 55.5% on an exam built so models fail it.19 Aug 2026 · Anthropic positioning and benchmark coverage (Claude Fable 5, August 2026)
- TRUE, BUT · Moved RulerTwo new benchmarks agree: the best AI models in the world clear fewer than half of a hard benchmark of real analyst tasks. Claude Fable 5 tops the frontier at 49.2%. The context is who built the tests, and who sells the fix.19 Aug 2026 · Samaya AI (FrontierFinance) and Vals AI (Finance Agent v2), August 2026
- TRUE, BUT · Lab Not FieldOpenAI and Anthropic are selling the same next step for AI agents: more of them. Claude Code now forks subagents by default, and Sol Ultra fans a problem across up to 64. Google Research ran the controlled test, and the answer is a split: more agents help work that breaks into independent pieces and hurt work that runs as one dependent chain, by up to 70%. Which one your task is decides whether the swarm is an upgrade or a tax.19 Aug 2026 · OpenAI (GPT-5.6 Sol Ultra) and Anthropic (Claude Code subagents), and the industry 'more agents are better' heuristic, August 2026
- TRUE, BUT · Moved RulerThe coding number everyone quotes says AI has nearly solved software engineering: Claude Opus 5 scores 96% on SWE-bench Verified. Move to the benchmark built to resist contamination and the top model sits at 80.3%, and GPT-5.6 Sol lands at 64.6%.19 Aug 2026 · Frontier-model coding marketing built on SWE-bench Verified (Anthropic, OpenAI and coverage, August 2026)
- CONTESTED93 percent of developers now use AI coding tools. Six independent studies converge on the same measured productivity gain: about 10 percent. In a randomized trial, experienced developers using frontier AI tools took 19 percent longer than those working without them. They thought they were 20 percent faster.18 Aug 2026 · AI tool vendors and companies citing AI productivity to justify spend