The check queue
The wire
Headlines from the sources we read, grouped by story. Checked stories link to the full evidence. Everything else is labelled unchecked. See a claim worth investigating? Ask the desk to check it.
From the checked collection
- BSMusk says SpaceX will have a GPT-6 level model in 2 to 3 months. He offered GPUs, not a benchmark.
- True, butGoogle's flagship voice model really is cheaper and better. The bargain one isn't.
- True, butOpenAI's mental health test, graded by OpenAI's model, scored clinicians below its AI.
- True, butGPT-6 Sol costs less. Its benchmark score barely moved.
- True, butAnthropic said Opus 5.5 runs 40% cheaper. The price list says 20%.
- True, butOpenAI shelved GPT-6.1 Astra over deception and acting without permission, and Sol 'doesn't have these problems.' OpenAI's own card shows Sol misrepresenting its work at 1.50%, against 0.51% for GPT-6 Astra.
- True, butMeta's Muse agent 'gave out a user's home address without permission.' The user later reviewed the logs with Meta and found he had clicked 'Allow Always.' That setting is the real story.
- BSNvidia says its new agent safety platform could have prevented the Hugging Face hack. No test against that hack is on the record.
- True, butAgents linked to OpenAI hit a UN data site about 16,500 times. The researcher who counted says he would not call it hacking.
- True, butOpenAI found a self-copying prompt injection in its own training runs. It says no impact was seen outside simulation.
- Holds upOpenAI did pause its most capable models. Its own report says the trigger was one agent that found a DNS gap.
- True, butRogue OpenAI agents 'meddled' with three US government sites. Two visits were public data; the one hack attempt did not succeed.
- ContestedAlbanese calls it a hack. OpenAI calls it a model that misbehaved.
- Holds upDaily AI users can still worry about it. Gallup's results support that headline.
- Holds upOxford's AI agents secretly colluded at blackjack. Catching them barely beats a coin flip.
- True, butThe 'first AI hack of a government' was three attempts. None appear to have worked.
- Holds upPerplexity's sandbox wall held in 108 runs. Its network fence did not, at first.
- Holds upEvilTokens used an old sign-in trick, amplified by AI.
- Too soonBloomberg reports AI overreliance in the Minab strike. The full probe is unreleased.
- Holds upGemini reached real companies during a test. The headline already said first for Google.
- Holds upClaude agents matched people's book tastes on 61% of pairs. Anthropic says a coin flip gets 50%.
- Too soonClaude flagged an enzyme system. Its function is still unknown.
- True, butEnveda doubled its valuation. Safety results do not prove weight-maintenance benefits.
- True, butAlibaba's cancer AI beats radiologists. It only reads abdominal CT scans.
- Holds upMeta's new audio glasses have no camera. Other privacy questions remain.
- True, butGoogle says even Google cannot read its new AI memory. The last named audit said safe from everyone except Google.
- True, butMeta called Muse 'safe and secure.' It launched with a 0-day.
- True, butMcDonald's is using AI to 'dynamically' price your burger. The Reuters report behind that line describes a per-store price recommender, and says it could not tie the Fresno Big Mac gap to it.
- True, butConsumer AI spending 'tripled to $40 billion' while users barely grew. Both numbers are Menlo Ventures estimates scaled up from a survey of 5,067 U.S. adults.
- True, but'DeepSeek $1B ARR' is a run rate from unnamed sources. The same reporting puts seven months of revenue at about $70.7 million.
- Too soonSnorkel reports a $375M revenue run rate. Its announcement does not show the calculation.
- True, butXiaomi's '$3M' top open model traces to one hedged tweet at $2.6M.
- True, butThe headline says US and Russia stripped human oversight from a UN AI weapons pact. The text still affirms human control.
- True, butThe Pentagon says an appeals court 'completely' validated its Anthropic blacklisting. It won on one of two designations.
- True, butGates did say AI could drive a billion deaths. The headlines cropped out the people with ill intent.
- BSMedicare's number two said AI prior-auth contractors don't earn more by denying care. Medicare's own page says they get a cut of care averted.
- True, butAnthropic 'files for a $2T IPO' with a $42B loss.
- True, butNvidia 'wants to put a watchdog chip next to every AI agent.' Per Nvidia's own developer blog, the watchdog is optional software on a BlueField-4 chip it announced in January.
- Holds upEmergence AI says it ran eight AI worlds of ten agents each.
Checked ruling + letter · In the queue being checked · Unchecked headline only, not ours · Lab a source's own blog, not press coverage of it
Loading the wire…