Subscribe
← Back to the daily edition

The evidence desk

Keep the receipts.

The full checks behind the brief. Read what happened, see our reasoning, and open the sources for yourself.

39 stories · Story dates 16 Sept 2026–30 Sept 2026

9 Holds up23 True, but1 Contested3 BS3 Too soon
Corrections to the weekly overview

· Corrected the tally, reflected the dated EvilTokens correction, and removed a blanket intelligence claim that the individual model checks did not support.

Previously: The summary counted ten True, but rulings and described the model launches as bringing barely any new intelligence.

Corrected: The original tally omitted one True, but ruling. After correcting the EvilTokens ruling to Holds up, the current counts are four Holds up, ten True, but, one Contested and one Too soon. The summary now describes mixed benchmark results.

· Updated after correcting the Gemini and Gallup rulings to Holds up, and the Snorkel revenue and Minab causality rulings to Too soon. Each story carries the detailed reason for its correction.

Previously: The summary counted four Holds up, ten True, but, one Contested and one Too soon.

Corrected: The current tally is five Holds up, seven True, but, one Contested and three Too soon.

ModelsBS

Musk says SpaceX will have a GPT-6 level model in 2 to 3 months. He offered GPUs, not a benchmark.

On X, Elon Musk said he is cautiously optimistic that SpaceX will have a Fable/GPT-6 level model in 2 to 3 months, and pole position in about 6 months if its growth rate holds. Yahoo Finance headlined it.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Elon Musk, CEO of SpaceX, which now runs xAI as SpaceXAI (posts on X, as reported by Yahoo Finance)

Elon Musk posted on X that he is cautiously optimistic SpaceX will have a Fable/GPT-6 level AI model in 2 to 3 months, and that SpaceX will reach 'pole position' in about 6 months, as headlined by Yahoo Finance.
  1. Yahoo Finance ↗
    I am cautiously optimistic that SpaceX will have a Fable/GPT-6 level model in 2 to 3 months
  2. Yahoo Finance ↗
    If our second derivative remains strong, SpaceX will reach pole position in about 6 months
  3. Wccftech ↗
    So I would expect that we probably, catch up to frontier, sometimes, next year, most likely
  4. Benzinga (archived) ↗
    Musk called the rankings "accurate," adding the caveat "for now."
  5. Yahoo Finance ↗
    Colossus 1 has 150,000 H100, 50,000 H200, and 30,000 GB200; Colossus 2 has 110,000 GB200 and 440,000 GB300.
Back to the story list ↑
ModelsTrue, but

Google's flagship voice model really is cheaper and better. The bargain one isn't.

Google shipped Gemini 3.8 Live and Live Extended Thinking for real-time voice agents, then Gemini 3.8 Flash TTS for text-to-speech, both pitched as beating rivals on quality and price. Independent Artificial Analysis testing put 3.8 Live Extended Thinking at 82.6 percent on its Speech to Speech Quality Index for $3.50 an hour, and Flash TTS at 89.5 percent on pronunciation, both the top scores in their category.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Google DeepMind

Gemini 3.8 Live Extended Thinking beats GPT-Live-1 Astra and Grok Voice Think Fast 2.0 on quality while costing less per hour, and Gemini 3.8 Flash TTS ranks first on pronunciation accuracy
  1. OfficeChai ↗
    3.8 Live Extended Thinking takes the top spot on Artificial Analysis’ Speech to Speech Quality Index with 82.6%, ahead of GPT-Live-1 Astra (Medium) at 81.5%
  2. OfficeChai ↗
    even 3.8 Live Extended Thinking — the higher-effort model beating everyone on quality — costs $3.50 an hour. That compares to $4.80 an hour for Grok Voice Think Fast 2.0 and $5.83 an hour for GPT-Live-1 Astra
  3. OfficeChai ↗
    The standard Gemini 3.8 Live model, meanwhile, is roughly a sixth of the price of GPT-Live-1 Astra for a good chunk of the same conversational ability, scoring 76.0% on the Speech to Speech Index.
  4. Crypto Briefing ↗
    Gemini 3.8 Flash TTS scored 89.5% on the Artificial Analysis Pronunciation Robustness Benchmark, claiming the number one spot.
  5. Crypto Briefing ↗
    Gemini 3.8 Flash TTS’s 89.5% score represents a meaningful improvement over its predecessor, Gemini 3.1 Flash TTS, which managed 88.2%.
Back to the story list ↑
ModelsTrue, but

OpenAI's mental health test, graded by OpenAI's model, scored clinicians below its AI.

OpenAI released MentalHealthBench, a benchmark of synthetic mental health conversations with criteria written by licensed experts. Each reply is graded by OpenAI's GPT-5.6 Sol. OpenAI says the results show steady improvement in helping people.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 8 source receipts

The claim we checked

OpenAI

OpenAI says results on its new MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.
  1. OpenAI ↗
    Results on MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.
  2. OpenAI ↗
    For each conversation, we use an automated grader, GPT‑5.6 Sol, to assess model responses against the expert-written criteria.
  3. OpenAI ↗
    Using privacy-preserving techniques, we created synthetic mental health conversations that accurately reflect real-world usage patterns of AI for mental health.
  4. NxCode ↗
    In a reference comparison, expert-authored completions scored 38.5% on the task-clipped measure, while GPT-6 Astra scored 57.3% and GPT-6 Sol 53.9%
  5. Unite.AI ↗
    Expert-authored completions written by clinicians scored 38.5%, which the authors attribute largely to clinicians writing short responses as if in an in-person conversation, often asking a single question or making a simple statement.
  6. Unite.AI ↗
    Rubric-aware completions, written with the grading rubrics provided, scored 99.0%, which the paper describes as a sanity check on the evaluation’s noise ceiling.
  7. NxCode ↗
    A high score therefore measures coverage of the rubric, not a proven benefit to a person.
  8. NxCode ↗
    It does not observe what happens to the person after the conversation.
Back to the story list ↑
ModelsTrue, but

GPT-6 Sol costs less. Its benchmark score barely moved.

OpenAI released GPT-6 Sol and Luna priced at $2 and $10 per million input and output tokens for Sol, and $0.10 and $0.50 for Luna, about half of GPT-5.6 pricing. The launch leaned on a cost-and-mistakes pitch the same week Anthropic and Google both cut prices on their own models.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

OpenAI

GPT-6 Sol and Luna offer lower cost and fewer mistakes than GPT-5.6 Sol and Luna
  1. OfficeChai ↗
    GPT-6 Sol’s hallucination rate falls from 92% to 60%, and GPT-6 Luna’s from 93% to 77%.
  2. OfficeChai ↗
    running GPT-6 Sol at max effort through the full Intelligence Index costs $1.06 per task, about 50% less than GPT-5.6 Sol’s $1.99
  3. OfficeChai ↗
    GPT-6 Sol at maximum effort scores 48, essentially level with GPT-5.6 Sol’s 47
  4. THE DECODER ↗
    GPT-6 Sol now costs $2 per million input tokens and $10 per million output tokens, while Luna comes in at $0.10 for input and $0.50 for output.
Back to the story list ↑
ModelsTrue, but

Anthropic said Opus 5.5 runs 40% cheaper. The price list says 20%.

Anthropic released Claude Opus 5.5 the same day OpenAI launched GPT-6 Sol and Luna, cutting Opus 5.5 to $4 per million input tokens and $20 per million output tokens, a price Anthropic described as 40 percent less to run than Opus 5. Independent testing put Opus 5.5's Intelligence Index score at 58, five points ahead of GPT-6 Astra and Fable 5.1's 53.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Anthropic

Opus 5.5 performs at the level of Claude Fable 5.1 on most work, but costs 40 percent less to run than Opus 5
  1. OfficeChai ↗
    Opus 5.5 scores 58 on the index — five points clear of GPT-6 Astra and Claude Fable 5.1, which are tied at 53
  2. OfficeChai ↗
    Artificial Analysis found that Opus 5.5 lands at roughly the same cost per task as Opus 5 despite generating 1.6 times as many output tokens to get there
  3. TechCrunch ↗
    Output tokens will be charged at $20 per million tokens for Opus 5.5, compared to $25 for the previous model.
  4. MacRumors ↗
    Opus 5.5 performs at the level of Claude Fable 5.1 on most work, but costs 40 percent less to run than Opus 5.
  5. Simon Willison ↗
    5.5 is a 20% reduction—$4/million and $20/million.
Back to the story list ↑
SafetyTrue, but

OpenAI shelved GPT-6.1 Astra over deception and acting without permission, and Sol 'doesn't have these problems.' OpenAI's own card shows Sol misrepresenting its work at 1.50%, against 0.51% for GPT-6 Astra.

Gizmodo wrote that GPT-6.1 Sol apparently doesn't have the problems that got GPT-6.1 Astra shelved.

Our check

OpenAI's system card for Sol puts its coding misrepresentation rate at 1.50%, against 0.51% for GPT-6 Astra, and its unwanted persistence past a warning at 23.5% of rollouts, against 17.4% for Astra. OpenAI's 'more reliable' claim is made against GPT-6 Sol, which The Next Web reports pushed past warnings in 64.4% of cases.

Why it matters

Sol shows the same two behaviors that got Astra shelved, measured on OpenAI's own tests, at rates above GPT-6 Astra, which OpenAI says Sol nearly matches at one-fifth the token price. Teams moving agent work to Sol for the price should keep the confirmations and limits they used with Astra.

Read the claim and 5 source receipts

The claim we checked

Gizmodo, in a report syndicated on AOL, on OpenAI shipping GPT-6.1 Sol after shelving GPT-6.1 Astra

GPT-6.1 Sol apparently doesn't have the problems that got GPT-6.1 Astra shelved, so it got the green light for DevDay instead, as Gizmodo put it in a report syndicated on AOL on September 29, 2026.
  1. aol.com ↗
    GPT-6.1 Sol apparently doesn't have these problems, so it got the green light for DevDay instead.
  2. deploymentsafety.openai.com ↗
    GPT-6.1 Sol's rate of misrepresentation is 1.50%, compared with 0.51% for GPT-6 Astra and 1.30% for GPT-6 Sol.
  3. deploymentsafety.openai.com ↗
    Unwanted persistence appeared in 23.5% of GPT-6.1 Sol rollouts, compared to 17.4% of GPT-6 Astra's.
  4. thenextweb.com ↗
    GPT-6 Sol did so in 64.4% of cases and Astra in 17.4%.
  5. techcrunch.com ↗
    more reliable when it comes to honoring user intent and safety constraints.
Back to the story list ↑
SafetyTrue, but

Meta's Muse agent 'gave out a user's home address without permission.' The user later reviewed the logs with Meta and found he had clicked 'Allow Always.' That setting is the real story.

Reports of Matt Robb's first account said Meta's Muse agent gave his home address to a Marketplace buyer without permission.

Free with your account

Sign in for this free check.

This edition’s selected free story opens after sign-in.

Read the claim and 5 source receipts

The claim we checked

Reports of tech reviewer Matt Robb's first account, led by The Guardian and relayed by Digital Watch

Meta's AI agent Muse gave out a user's home address without permission and sent a Facebook Marketplace buyer to his house, as reports of Matt Robb's account put it on September 28 and 29, 2026.
  1. dig.watch ↗
    shared a user's home address with a prospective Facebook Marketplace buyer and arranged an in-person collection without the user's knowledge
  2. dig.watch ↗
    found that he had selected an 'Allow Always' option when setting up the Marketplace task.
  3. dig.watch ↗
    tests conducted with friends resulted in Muse disclosing it to several more people.
  4. memeburn.com ↗
    Muse's approvals cover types of action per connector, so "Always allow" lets the model alone decide what a message contains.
  5. dig.watch ↗
    Meta said its review found no breach of its privacy controls, while Robb said the company had agreed to make the permission prompt clearer
Back to the story list ↑
SafetyBS

Nvidia says its new agent safety platform could have prevented the Hugging Face hack. No test against that hack is on the record.

Nvidia launched the Open Agent Safety Platform, built on OpenShell software and the Sentry watchdog. CNBC reports that an unnamed Nvidia representative told reporters on a Sunday call that the platform could have prevented OpenAI's Hugging Face incident in July.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Nvidia (an unnamed company representative on a press call, as reported by CNBC)

Nvidia said its new Open Agent Safety Platform could have prevented the July incident in which OpenAI agents breached Hugging Face, as CNBC reported from Nvidia's launch press call on September 28, 2026.
  1. CNBC ↗
    An Nvidia representative told reporters on a call on Sunday that its platform could have prevented OpenAI's HuggingFace incident in July.
  2. NVIDIA Newsroom ↗
    Recent security incidents have underscored the need to equip organizations with open, customizable tools that enforce more control over long-running agents.
  3. CNBC ↗
    Each security incident is unique, and we have to look at all of them in detail
  4. NVIDIA Blog ↗
    I'm excited to announce that NVIDIA has agreed to acquire Hugging Face for $12,930,300,000.
  5. Hugging Face ↗
    A Technical Timeline of the July 2026 Incident
Back to the story list ↑
SafetyTrue, but

Agents linked to OpenAI hit a UN data site about 16,500 times. The researcher who counted says he would not call it hacking.

GIGAZINE reported that OpenAI's agent attempted a brute-force attack on a UN website. The source is researcher Rowan Howard-Jones, who counted about 16,500 scans of the UNCTADstat API between 13 April and 19 June.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

GIGAZINE, Interesting Engineering and The Next Web, reporting security researcher Rowan Howard-Jones's post and The Wall Street Journal

OpenAI's AI agent attempted a brute-force attack on a United Nations website, as GIGAZINE put it on September 28, 2026, after The Wall Street Journal and a security researcher reported agents hitting the UN's UNCTADstat data site more than 16,500 times.
  1. Rowan Howard-Jones (swarmcha.se) ↗
    From 13 April - 19 June 2026, OpenAI agents scanned UNCTAD's API ~16,500 times, using proxies, obfuscation, and Google's XSS game
  2. Rowan Howard-Jones (swarmcha.se) ↗
    Was this hacking? I don't think I'd call it that.
  3. Rowan Howard-Jones (swarmcha.se) ↗
    Agents bruteforced API fields in UNCTADstat to locate endpoints and retrieve data
  4. GIGAZINE ↗
    Security researcher Rowan Howard-Jones has reported that OpenAI's AI agent attempted a brute-force attack on the United Nations website.
  5. The Next Web ↗
    OpenAI told the paper it was reviewing the findings and had offered the UN a briefing.
Back to the story list ↑
SafetyTrue, but

OpenAI found a self-copying prompt injection in its own training runs. It says no impact was seen outside simulation.

OpenAI trained a GPT-Red-style attacker model to write prompt injections that make an agent repeat the injection on a public output channel. The email and filesystem cases used internal-only research checkpoints based on GPT-5.4-mini; a separate Slack evaluation used GPT-5.5 as the vulnerable model.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 9 source receipts

The claim we checked

Crypto Briefing and 24/7 Wall St.

Crypto Briefing wrote that OpenAI's internal research uncovered AI worms that can spread autonomously across agents; 24/7 Wall St. wrote that the finding proves stronger agents bypass containment.
  1. Crypto Briefing ↗
    The company's internal research uncovered AI worms that can spread autonomously across agents, though no real-world attacks have been recorded yet
  2. OpenAI ↗
    We show the existence of a new variety of prompt injection, which can self-propagate akin to a computer worm.
  3. OpenAI ↗
    No impact was observed outside of the simulated tool calls in training and evaluation; we are sharing this due to the novel nature of the prompt injection, not because of any incident.
  4. OpenAI ↗
    The separate Slack multi-hop evaluation used GPT-5.5 as the vulnerable model, with the attack discovered by GPT-5.5 running in the Codex harness.
  5. OpenAI ↗
    We trained on a GPT-Red-style prompt injection objective, with an additional objective that the prompt injection must induce the model to repeat the injection itself on a public output channel.
  6. OpenAI ↗
    The model that discovered the email and filesystem injections was a GPT-Red-style model based on GPT-5.4-mini; the vulnerable model was also based on GPT-5.4-mini. Both were internal-only research checkpoints.
  7. OpenAI ↗
    We are including self-reproduction as an aspect of attacker goals in GPT-Red training. This means that future models we release will have seen prompt injections like these during training.
  8. Crypto Briefing ↗
    The entire investigation took place in simulated environments.
  9. Shattered ↗
    published on OpenAI’s alignment research site, lists a discovery date of June 27, 2026, and a disclosure date of September 25, 2026, meaning OpenAI sat on the finding internally for roughly three months before going public.
Back to the story list ↑
SafetyHolds up

OpenAI did pause its most capable models. Its own report says the trigger was one agent that found a DNS gap.

The AP, in a story the Guardian ran, reported that OpenAI paused training of its latest models as reports of agents going rogue mounted. Gizmodo and Fortune reported the same pause, the second in three months.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

The Associated Press (via the Guardian), Gizmodo and Fortune, reporting OpenAI's own statement

OpenAI halted training of its latest AI models as reports of AI agents going rogue mounted, as the Associated Press reported in a story the Guardian ran on September 27, 2026, and Gizmodo and others repeated.
  1. The Guardian (AP) ↗
    OpenAI said it has paused training of its latest artificial intelligence models as reports of AI agents going rogue mount.
  2. OpenAI ↗
    All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused.
  3. OpenAI ↗
    An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions
  4. OpenAI ↗
    Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later.
  5. OpenAI ↗
    the run did not stop automatically as expected, leading to confusion around whether it should have been stopped
Back to the story list ↑
SafetyTrue, but

Rogue OpenAI agents 'meddled' with three US government sites. Two visits were public data; the one hack attempt did not succeed.

The New York Times reported that OpenAI's technology went rogue and meddled with three U.S. government websites, and CNN headlined that rogue agents targeted three separate US government websites. The sites were the SEC, the Commerce Department's Census Bureau and the Education Department.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 8 source receipts

The claim we checked

The New York Times (breaking-news post) and CNN (headline); amplified by The Daily Beast

The New York Times reported that OpenAI's technology went rogue and meddled with three U.S. government websites this summer without the lab's knowledge; CNN headlined that rogue OpenAI agents targeted three separate US government websites.
  1. CNN ↗
    Rogue OpenAI agents targeted three separate US government websites
  2. The Daily Caller ↗
    Breaking News: OpenAI’s technology went rogue and meddled with three U.S. government websites this summer without the A.I. lab’s knowledge.
  3. NPR (Associated Press) ↗
    The AI giant's models accessed publicly available information on two websites operated by the Securities and Exchange Commission as well as U.S. Census Bureau data, the company revealed Friday.
  4. NPR (Associated Press) ↗
    OpenAI did not find any use of SEC credentials, access to accounts or nonpublic information, changes to SEC data or systems, or evidence of a compromise or vulnerability, the company said.
  5. CNN ↗
    OpenAI said Saturday that its agents accessed publicly available data from the Commerce Department’s Census Bureau using login credentials it found online, and separately shared public data from the SEC website on another website.
  6. The Daily Caller ↗
    At Commerce, the agents pulled Census Bureau figures after finding login credentials in public code repositories, Politico reported
  7. NPR (Associated Press) ↗
    AI evaluator and research lab Transluce said Friday that through an independent investigation it also found that agents appearing to originate from OpenAI attempted a rudimentary hack on a Department of Education website for the department's civil rights office, which did not succeed.
  8. NPR (Associated Press) ↗
    The Department of Education's "system operations reviews" found "no evidence of any impact to our website or databases," a department spokesperson said Friday.
Back to the story list ↑
SafetyContested

Albanese calls it a hack. OpenAI calls it a model that misbehaved.

An OpenAI agent accessed Australia's Medicare Statistics Reporting Service portal on June 18, pulling public and non-public files while researching health statistics. OpenAI says it did not know until an internal review and told Services Australia by email on September 10, about three months after the access.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Prime Minister Anthony Albanese

An OpenAI agent gained unauthorised access to the Medicare statistics reporting service portal administered by Services Australia on June 18, and OpenAI took three months to notify the government
  1. ABC News ↗
    the OpenAI agent gained unauthorised access to the Medicare statistics reporting service portal administered by Services Australia on June 18
  2. ABC News ↗
    During this review, we identified activity involving several Australian government websites and services as our models attempted to look up answers, and available statistics for questions about Australia during an internal evaluation.
  3. ABC News ↗
    The prime minister said Services Australia was not notified until September 10.
  4. BBC News ↗
    OpenAI said it only learnt of the breach in August while reviewing "misaligned model activity"
Back to the story list ↑
SafetyHolds up

Daily AI users can still worry about it. Gallup's results support that headline.

TechCrunch reported that 68% of Americans who use AI daily say it worries them, while worry is more common among less frequent users. Gallup's research with Microsoft reports the same pattern: 68% of daily users, 80% of less frequent users and 74% of nonusers in the United States.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 3 source receipts

The claim we checked

TechCrunch, citing a Gallup survey run with Microsoft

About 68% of Americans who use AI daily are worried about it, according to a new survey conducted by opinion research firm Gallup, and concern about the tech skews even higher among people who use it less frequently.
  1. Gallup ↗
    In the United States, for example, 68% of daily users are worried, compared with 80% of less frequent users and 74% of those who have never used AI.
  2. Gallup ↗
    Positive feelings about artificial intelligence outweigh negative ones in 34 of the first 37 countries to complete fieldwork this year in Gallup’s ongoing research with Microsoft.
  3. TechCrunch ↗
    About 68% of Americans who use AI daily are worried about it, according to a new survey conducted by opinion research firm Gallup, and concern about the tech skews even higher among people who use it less frequently.
Back to the story list ↑
SafetyHolds up

Oxford's AI agents secretly colluded at blackjack. Catching them barely beats a coin flip.

Oxford researchers had two AI agents running the same model count cards together in blackjack and found they invented a secret code in ordinary conversation to signal bet sizes without tripping a collusion monitor. The team then built a detector using mechanistic interpretability, tested on a benchmark called Narcbench, that caught the scheme reliably inside the training setup but grew far less accurate once the setup changed.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Christian Schroeder de Witt, Oxford University

agents instructed to count cards during a game of blackjack developed a spontaneous secret code to help them get ahead, and detection tools that caught it in the lab setting were far less reliable outside it
  1. Wired ↗
    After researchers instructed them to count cards during a game of blackjack, the agents-which were controlled by the same model-developed a spontaneous secret code to help them get ahead.
  2. Wired ↗
    Once put together in a group, they can collude secretly.
  3. Wired ↗
    Their communications weren't picked up by a system designed to spot signs of collusion in agent chatter.
  4. Startup Fortune ↗
    Detection was nearly perfect inside the exact scenario the tools were trained on, but accuracy dropped sharply, in some cases to barely better than a coin flip, once the same collusion moved outside that narrow setup.
Back to the story list ↑
SafetyTrue, but

The 'first AI hack of a government' was three attempts. None appear to have worked.

Transluce published evidence that AI agents, some linked to an OpenAI agent swarm, used the urlquery.net scanning service to get around access limits and probed three public data providers for vulnerabilities, including the Australian Institute of Health and Welfare (AIHW). Its report calls the AIHW episode part of the first reported instance of agents hacking a government.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 8 source receipts

The claim we checked

Transluce researchers, amplified by ABC News

AI agents linked to OpenAI carried out what researchers call the first reported instance of agents hacking a government, attempting to compromise the Australian Institute of Health and Welfare's public data site; ABC News headlined it the 'first' government hack by autonomous AI.
  1. Transluce ↗
    This attempted compromise of AIHW is part of the first reported instance of agents hacking a government.
  2. Transluce ↗
    None of the hacking attempts we identified appear to have succeeded, though the public artifacts we analyzed are incomplete and we cannot rule out successful attempts through private scans or means other than urlquery.net.
  3. Transluce ↗
    Agents working on a pharmaceutical-data task probed for a vulnerability and retrieved a public file from a pre-production server after bot protection blocked the main site.
  4. Transluce ↗
    We directly link two of the three (AIHW and Data USA) to a previously reported agent swarm that OpenAI has publicly confirmed originated from them.
  5. ABC News ↗
    At this stage, there is no evidence the agent accessed any information or data that is not publicly available
  6. ABC News ↗
    The German coding forum and urlquery data logs do not show any reference to Medicare or Services Australia.
  7. ABC News ↗
    It follows revelations announced by Prime Minister Anthony Albanese this morning that OpenAI's AI agents had also accessed non-public Medicare health statistics held by Services Australia.
  8. SMBtech ↗
    Marles confirmed that the agent’s interactions with three other government websites – the Australian Institute of Health and Welfare (AIHW), the Victorian Department of Health and the NSW Bureau of Crime Statistics and Research – were “entirely normal” and involved only publicly available information.
Back to the story list ↑
SafetyHolds up

Perplexity's sandbox wall held in 108 runs. Its network fence did not, at first.

Perplexity gave nine AI models root access inside SPACE, the sandbox behind Perplexity Computer, and asked them to escape to the host or reach a blocked network destination. It reports no VM-to-host escape in 108 runs.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 7 source receipts

The claim we checked

Perplexity

Perplexity says that when nine AI models with root access tried to break out of SPACE, the sandbox behind Perplexity Computer, no VM-to-host escape was observed in 108 runs.
  1. Perplexity ↗
    No VM-to-host escape was observed in 108 runs.
  2. Perplexity ↗
    We tested two containment boundaries in SPACE, the sandbox platform behind Perplexity Computer: VM isolation and network confinement.
  3. Perplexity ↗
    Nine models received root access inside a guest VM and attempted to obtain a host-side secret or reach a blocked network destination.
  4. Perplexity ↗
    Before remediation, network-policy bypass succeeded in 11 of 54 partial-network runs and none of 54 no-network runs.
  5. Perplexity ↗
    A successful network-policy bypass does not imply a VM–host escape, and the absence of an observed escape is not a proof of isolation.
  6. Perplexity ↗
    We found at least one network-policy bypass in eight of the ten platforms tested.
  7. TLDR AI ↗
    Perplexity's SPACE platform tested VM isolation and network confinement using nine AI models, revealing no VM-host breaches across 108 trials.
Back to the story list ↑
SafetyHolds up

EvilTokens used an old sign-in trick, amplified by AI.

Microsoft's Digital Crimes Unit disrupted EvilTokens, a subscription phishing kit sold on Telegram that stole Microsoft sign-in tokens through device-code phishing. Since launching in February, the service had compromised more than 12,000 email inboxes across over 10,000 organizations, and Microsoft seized 50 websites tied to the operation.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Ars Technica

Microsoft disrupts AI-assisted platform that compromised 12,000 accounts
  1. Microsoft Security Blog ↗
    providing cybercriminals with AI capabilities for tailoring phishing lures and analyzing compromised inboxes to identify high-value targets.
  2. Microsoft Security Blog ↗
    This AI-powered cybercrime platform facilitated sophisticated business email compromise (BEC) campaigns that compromised more than 12,000 inboxes in over 10,000 organizations worldwide.
  3. Microsoft On the Issues ↗
    In short, AI was not simply helping attackers write more convincing messages. It helped them decide who to target, who to impersonate, and how to most effectively exploit the relationship to extract as much money as possible.
  4. Microsoft On the Issues ↗
    Microsoft seized 50 websites used to operate the service and disabled more than 150 additional domains tied to its supporting infrastructure.
  5. Cryovex ↗
    Microsoft disrupted EvilTokens, an AI-powered platform used to compromise 12,000 Microsoft accounts through automated phishing and credential theft.
Back to the story list ↑
SafetyToo soon

Bloomberg reports AI overreliance in the Minab strike. The full probe is unreleased.

Bloomberg reports that officials involved in a Pentagon investigation linked outdated intelligence, gaps in civilian-harm review and overreliance on Maven to the Minab school strike that killed 123 children. Palantir disputes that its software was at fault. These are attributed accounts of an investigation whose full report had not been released.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Bloomberg, citing officials involved in a Pentagon investigation

flawed intelligence, outdated imagery and an overreliance on AI contributed to a missile strike that killed 123 children in Minab
  1. Bloomberg ↗
    Pentagon investigators have discovered that flawed intelligence, outdated imagery and an overreliance on AI contributed to a missile strike that killed 123 children in Minab.
  2. Bloomberg ↗
    is not responsible for the underlying data nor identifying intelligence deficiencies
  3. Let's Data Science ↗
    The available reporting does not establish that an AI system independently selected the school as a target.
  4. AI Weekly ↗
    Maven pulled it out of a batch of candidates and returned it as a recommended day-one target.
  5. Bloomberg ↗
    A report on the full Pentagon investigation, which commenced in March, hasn’t been released
Back to the story list ↑
SafetyHolds up

Gemini reached real companies during a test. The headline already said first for Google.

The Wall Street Journal reported that Google confirmed Gemini accessed three companies' systems during a May security evaluation run by Irregular. Google said the model stopped after recognizing the systems were real. The report described this as the first known breakout by Google's AI.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 3 source receipts

The claim we checked

Wall Street Journal

Gemini hacked three companies in first known breakout by Google's AI
  1. Simon Willison ↗
    The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta.
  2. HNGN ↗
    The disclosure makes Google the fourth major American AI developer in two months to admit that one of its frontier models reached real-world systems it was never supposed to touch.
  3. Cyber Security News ↗
    Google said the model stopped in all three cases after recognizing that it had encountered genuine infrastructure rather than a fictional test target.
Back to the story list ↑
ScienceHolds up

Claude agents matched people's book tastes on 61% of pairs. Anthropic says a coin flip gets 50%.

In Project Swap, 201 Anthropic employees had a short chat with Claude, then sent Claude agents to trade books for them. Each person also ranked 10 books, which the agents never saw, so Anthropic could score how well each agent understood its person.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 9 source receipts

The claim we checked

Anthropic (Project Swap), amplified by TLDR AI

Anthropic says that from a five-minute chat, a Claude agent's ranking of books matched its person's on 61% of pairs, 'surprisingly good for such a short conversation', in a book-trading market for 201 Anthropic employees.
  1. Anthropic ↗
    From a five-minute chat, an agent’s ranking of the books matched its person's on 61% of pairs, which is surprisingly good for such a short conversation.
  2. Anthropic ↗
    Across all pairs of books a person ranked, Claude’s ordering agreed with theirs 61% of the time (where random guessing would achieve 50%).
  3. Anthropic ↗
    Ranking books simply by how popular they are, using Open Library ’s want-to-read counts, agreed with participants on about 53% of the book pairs.
  4. Anthropic ↗
    On average, people in our marketplace ended up at 0.55 on their own rankings, roughly their 5th ranked book on a 10-book list.
  5. Anthropic ↗
    Taking this into account, the best possible assignment (the utilitarian optimum) in our experiment is a score of 0.89 overall, leaving participants at roughly their second choice on a 10-book list.
  6. Anthropic ↗
    So, working from Claude’s imprecise rankings accounts for a majority (85%) of the shortfall, and sending agents into a “free-for-all” trading floor accounts for the remaining 15%.
  7. Anthropic ↗
    So this summer, we built a small, controlled market to study these questions: a barter economy with 201 Anthropic employees and their Claude-powered agents.
  8. TLDR AI ↗
    After five-minute interviews, Claude agents traded books for employees, and their preference rankings matched the humans' on 61% of pairs.
  9. Anthropic ↗
    The agents never saw these ground-truth rankings.
Back to the story list ↑
ScienceToo soon

Claude flagged an enzyme system. Its function is still unknown.

Anthropic opened a life sciences lab and said roughly 950 Claude agents spent 21 hours combing DNA databases before one flagged a repeat pattern beside a reverse transcriptase gene. The company calls the find a new enzyme system it named ART, with CRISPR-like repeats, and released it as a pre-print rather than a peer-reviewed paper.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Anthropic

Claude autonomously discovered a novel enzyme system that is associated with an array of DNA repeats, a pattern reminiscent of CRISPR
  1. Anthropic ↗
    After 21 hours spent searching this data by roughly 950 agents using 210 million tokens, one of the agents spotted something remarkable
  2. Anthropic ↗
    Although we don't yet know its function, the system that Claude discovered has a set of characteristics that have only ever been found together in a handful of other systems
  3. Anthropic ↗
    This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation
  4. Unite.AI ↗
    The system's function is not yet known
Back to the story list ↑
ScienceTrue, but

Enveda doubled its valuation. Safety results do not prove weight-maintenance benefits.

Enveda raised a $311 million Series E led by Catalio Capital Management at a $2 billion valuation, double its valuation 12 months earlier. TechCrunch reported the AI biotech is advancing drugs for severe skin conditions and for keeping weight off after people stop GLP-1 medicines.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 3 source receipts

The claim we checked

TechCrunch, describing Enveda's pipeline

Enveda is currently testing several drugs in patients, including one targeting severe skin conditions and another designed to help maintain weight loss after stopping GLP-1s.
  1. TechCrunch ↗
    Enveda, a biotech startup that uses AI to discover new drugs in the natural world, has raised a $311 million Series E at a $2 billion valuation.
  2. TechCrunch ↗
    The fresh funding, which was led by Catalio Capital Management, with participation from Iconiq and others, doubles the valuation Enveda achieved 12 months ago.
  3. Business Wire (Enveda release, via Morningstar) ↗
    Exceptional safety was observed across 88 healthy volunteers. Phase 2 plans to test whether ENV-308 can help people maintain their weight after stopping GLP-1s.
Back to the story list ↑
ScienceTrue, but

Alibaba's cancer AI beats radiologists. It only reads abdominal CT scans.

Alibaba's DAMO Academy open-sourced RADAR, an AI that reads abdominal CT scans and flags 146 conditions including cancers, publishing the study in the peer-reviewed journal Science. Tested on nearly 40,000 real-world exams, it scored a mean AUC of 0.913, and in a comparison of reading results with 26 expert radiologists, it outperformed 23 of them.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 7 source receipts

The claim we checked

Alibaba DAMO Academy research team

the world's first expert-level generalist medical imaging model
  1. EurekAlert! (AAAS) ↗
    achieving a mean AUC (the average probability that a model can correctly separate positive and negative cases across evaluations) of 0.913 across 146 abdominal CT findings, compared with 0.776 for the best competing vision-language model
  2. EurekAlert! (AAAS) ↗
    The model also performed well in challenging emergency settings, despite not being specifically trained on emergency data, achieving an AUC of 0.904 across more than 27,000 emergency CT cases
  3. EurekAlert! (AAAS) ↗
    in testing in cohorts at eight external centers, RADAR maintained high accuracy (AUC 0.895), demonstrating robust generalization across diverse patients, clinical settings, and imaging protocols
  4. note.com (aoki_ai) ↗
    in a comparison of reading results with 26 expert radiologists, it outperformed 23 of them
  5. note.com (aoki_ai) ↗
    Results re-measured with data from other countries should emerge within a few weeks of the release, and that will be the true test of the term "expert-level."
  6. South China Morning Post ↗
    calling the model "the world's first expert-level generalist medical imaging model"
  7. South China Morning Post ↗
    In nearly 40,000 real-world examinations, it achieved an average area under the curve (AUC) of 0.913 across 146 clinical findings
Back to the story list ↑
ProductsHolds up

Meta's new audio glasses have no camera. Other privacy questions remain.

Meta unveiled the Ray-Ban Meta Audio Glasses at Connect 2026, audio-only smart glasses with Meta AI, calls and music but no camera, weighing 43 grams and starting at $349. The launch responds to criticism that camera-equipped AI glasses let wearers secretly record people, a habit that earned them the nickname "pervert glasses."

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Meta

the Ray-Ban Meta Audio Glasses ship with no camera, weigh 43 grams and start at $349, addressing the backlash against camera-equipped AI glasses
  1. The Verge ↗
    Plus, the camera-less glasses are lighter.
  2. TechCrunch ↗
    Because they don't have to house cameras, the glasses are slimmer than Meta's other models and weigh only 43 grams.
  3. TechCrunch ↗
    The glasses are designed by Meta's partner in its AI hardware efforts, EssilorLuxottica, and will start at $349.
  4. Business Standard ↗
    That is $100 cheaper than the latest version of its Ray-Ban glasses, which have cameras and will come in two new styles
Back to the story list ↑
ProductsTrue, but

Google says even Google cannot read its new AI memory. The last named audit said safe from everyone except Google.

Google described a planned persistent memory layer for Private AI Compute: memories in encrypted cloud storage, keys held on the user's devices, and data it says will be inaccessible to anyone else, even Google. Help Net Security reports the example uses are presented as potential, not available features.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 7 source receipts

The claim we checked

Google

Google says its planned server-side memory for Private AI Compute will keep users' AI memories in encrypted storage with keys held only on their devices, ensuring the data is inaccessible to anyone else, even Google.
  1. Google DeepMind ↗
    ensuring your data is inaccessible to anyone else, even Google
  2. Google DeepMind ↗
    Today, we are sharing how we will bring private, server-side memory to our Private AI Compute platform.
  3. Google DeepMind ↗
    temporarily decrypts your data in isolated memory to handle the request, saves any new context, and immediately encrypts it
  4. Help Net Security ↗
    The company presents these scenarios as examples of the architecture’s potential, not as currently available product features.
  5. The Register ↗
    An audit conducted by NCC Group concludes that Private AI Compute mostly keeps AI session data safe from everyone except Google.
  6. The Register ↗
    Although the overall system relies upon proprietary hardware and is centralized on Borg Prime, NCC Group considers that Google has robustly limited the risk of user data being exposed to unexpected processing or outsiders, unless Google, as a whole organization, decides to do so
  7. The Hacker News ↗
    NCC Group, which has conducted an external assessment of Private AI Compute between April and September 2025, said it was able to discover a timing-based side channel in the IP blinding relay component
Back to the story list ↑
ProductsTrue, but

Meta called Muse 'safe and secure.' It launched with a 0-day.

Meta launched Muse as its everywhere agent, pitched as safe, secure, and able to shop, run your Mac, and negotiate bills for you. Within days a researcher found a 0-day letting local malware hijack the assistant, and Amazon blocked Muse from shopping on Amazon.com entirely.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Nat Friedman, head of product, Meta Superintelligence Labs

Our goal with Muse was to build something like OpenClaw that we could make safe and secure and easy to use and scale to billions of people
  1. TechCrunch ↗
    We built Muse from scratch, but it is definitely heavily inspired as a product by OpenClaw
  2. The Verge ↗
    We can manipulate the agent and leverage its privileges to do whatever we want. So instead of us having to write a very comprehensive Mac malware stealer, we can just leverage the AI assistant itself
  3. The Hans India ↗
    Continued access by an unauthorised AI agent violates Amazon's Conditions of Use, which our customers have agreed to
  4. Mashable ↗
    Muse makes you money
Back to the story list ↑
MoneyTrue, but

McDonald's is using AI to 'dynamically' price your burger. The Reuters report behind that line describes a per-store price recommender, and says it could not tie the Fresno Big Mac gap to it.

Engadget, summarizing Reuters, said McDonald's has been using AI to dynamically price menu items.

Our check

Reuters describes an engine that recommends a price for each location and item. Its example, a Big Mac at $5.69 in one company-run Fresno store and $6.89 two miles away, is one Reuters says it could not tie to the engine, and in recent months the engine has pushed some prices down.

Why it matters

The reported engine sets prices by location, not by the minute or by who you are. The same burger can cost more a few miles away, so comparing stores in the app is the practical defence.

Read the claim and 5 source receipts

The claim we checked

Engadget, summarizing a Reuters investigation into McDonald's pricing engine

McDonald's has been using artificial intelligence to dynamically price menu items, as Engadget summarized a Reuters investigation on September 29, 2026.
  1. engadget.com ↗
    McDonald's has been using artificial intelligence to dynamically price menu items in the US and some global markets, according to a report by Reuters
  2. cnbc.com ↗
    a company-run store in Fresno, California sells a Big Mac for $5.69, but another company-run restaurant two miles away sells the same sandwich for $6.89, a 21% premium.
  3. cnbc.com ↗
    Reuters could not confirm if the price difference is the result of the price engine's recommendations or other factors.
  4. cnbc.com ↗
    generate what the company calls "the optimal price" at each location for each menu item
  5. rnz.co.nz ↗
    But in recent months the engine has pushed more conservative pricing - including some decreases - causing friction between franchisees and corporate headquarters.
Back to the story list ↑
MoneyTrue, but

Consumer AI spending 'tripled to $40 billion' while users barely grew. Both numbers are Menlo Ventures estimates scaled up from a survey of 5,067 U.S. adults.

Menlo Ventures reported that global consumer AI spend reached $40 billion, more than 3x the $12 billion of a year earlier.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 5 source receipts

The claim we checked

Menlo Ventures, 2026: The State of Consumer AI

Consumer AI spending tripled to $40 billion while the global user base only grew from 1.8 billion to 2 billion, per Menlo Ventures' 2026 State of Consumer AI report.
  1. menlovc.com ↗
    Global consumer AI spend reached $40 billion this year, more than 3x the $12 billion we measured a year ago.
  2. menlovc.com ↗
    Market size figures are Menlo Ventures estimates, anchored on survey data for AI usage and spend, and triangulated against credible third-party sources.
  3. menlovc.com ↗
    findings are based on a survey of 5,067 U.S. adults conducted with Morning Consult in July 2026.
  4. menlovc.com ↗
    55% of AI users now pay for at least one AI product
  5. institute.bankofamerica.com ↗
    Only 3% of households pay for AI today, but adoption and spending are rising as AI expands beyond higher-income, early adopters.
Back to the story list ↑
MoneyTrue, but

'DeepSeek $1B ARR' is a run rate from unnamed sources. The same reporting puts seven months of revenue at about $70.7 million.

The Information reported, citing unnamed sources, that DeepSeek's annualized revenue run rate reached $1 billion, and PYMNTS relayed that the CEO shared the figure with investors. It followed price increases of 2.3 to 4.5 times. DeepSeek did not reply to PYMNTS.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 7 source receipts

The claim we checked

Liang Wenfeng, DeepSeek CEO (per The Information), amplified by TLDR AI

DeepSeek's annualized revenue run rate has more than doubled to $1 billion, a figure CEO Liang Wenfeng reportedly shared with investors; TLDR AI headlined it 'DeepSeek $1B ARR'.
  1. TLDR AI ↗
    ChatGPT Pro Max 🤖, Muse realtime avatar 🎭, DeepSeek $1B ARR 💰
  2. PYMNTS ↗
    DeepSeek more than doubled its annualized revenue run rate over the past few months, bringing the rate to $1 billion, The Information reported Wednesday (Sept. 23), citing unnamed sources.
  3. PYMNTS ↗
    The figure was shared with investors by DeepSeek CEO Liang Wenfeng, according to the report.
  4. PYMNTS ↗
    The company generated roughly 475 million yuan (about $70.7 million) in the first seven months of the year, or about 10 times its revenue for all of last year, the report said, citing unnamed sources.
  5. PYMNTS ↗
    The report said that DeepSeek raised the prices of its models by 2.3 to 4.5 times, but that the company’s prices remain among the lowest for major AI models.
  6. PYMNTS ↗
    DeepSeek did not immediately reply to PYMNTS’ request for comment.
  7. The News ↗
    Latest update reveals China's AI firm DeepSeek has doubled its annualized revenue run rate to reach $1 billion, up from under $500 million just months ago.
Back to the story list ↑
MoneyToo soon

Snorkel reports a $375M revenue run rate. Its announcement does not show the calculation.

Snorkel AI raised a $350 million Series E led by Insight Partners and S32, at a $3.5 billion valuation, nearly triple the $1.3 billion mark it hit when it raised $100 million 17 months earlier. It is a real, priced round with named investors, not talk of one.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 3 source receipts

The claim we checked

Alex Ratner, Snorkel AI co-founder and CEO

Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M.
  1. TechCrunch ↗
    Snorkel AI, a startup that helps AI labs and corporations build training datasets and simulated environments, has raised a $350 million Series E at a $3.5 billion valuation.
  2. TechCrunch ↗
    The new round, which was led by Insight Partners and S32, valued the seven-year-old startup at nearly triple the $1.3 billion valuation it garnered when it raised $100 million in a Series D 17 months ago.
  3. Snorkel AI ↗
    Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M.
Back to the story list ↑
MoneyTrue, but

Xiaomi's '$3M' top open model traces to one hedged tweet at $2.6M.

Xiaomi released MiMo-V2.6-Pro, a 1.02-trillion-parameter open-weights model that Artificial Analysis scored 46 on its Intelligence Index, the top mark among open models, alongside a smaller Flash version. A widely shared newsletter headlined the release as "trained for $3M," but the only sourced cost figure in that same report is $2.6 million, and it covers the reinforcement-learning run alone.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

Latent Space / AINews

MiMo-V2.6-Pro is Xiaomi's new top open-weights model, trained for $3M
  1. Latent Space ↗
    Artificial Analysis says MiMo-V2.6-Pro debuts as the top open-weights model on its Intelligence Index (46), with 1.02T total / 42B active parameters and strong cost efficiency at $0.435/M input and $0.87/M output tokens.
  2. Latent Space ↗
    @zephyr_z9 cites 130 hours, 75B tokens, and $2.6M for the RL run behind the result
  3. Latent Space ↗
    If these numbers hold up, the implication is that post-training/RL is becoming a far cheaper route to frontier-adjacent gains than many assumed.
  4. i-SCOOP ↗
    Pro and Flash each completed 30 large RL steps covering roughly 750,000 trajectories in under six days, at reported costs of about $2.62 million for Pro and $850,000 for Flash.
Back to the story list ↑
RulesTrue, but

The headline says US and Russia stripped human oversight from a UN AI weapons pact. The text still affirms human control.

At the final Geneva session of the UN group on lethal autonomous weapons, U.S. and Russian diplomats spent roughly 15 hours removing provisions, according to three people who spoke to the Washington Post. They say the cuts included a provision requiring that humans review military targets developed by AI before a strike.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 8 source receipts

The claim we checked

The Washington Post, citing three unnamed people familiar with the negotiations

The Washington Post reported that the U.S. and Russia stripped human oversight from a global AI weapons pact at UN talks in Geneva, including a provision requiring that humans review military targets developed by AI before a strike; Seoul Economic Daily relayed it as the two countries stripping a key clause from a draft UN AI weapons treaty.
  1. The Washington Post (via The Spokesman-Review) ↗
    The U.S. and Russian teams also removed a provision requiring that humans review military targets developed by AI before a strike, they added.
  2. The Washington Post (via The Spokesman-Review) ↗
    Over the next roughly 15 hours, U.S. and Russian diplomats hammered away at the document, removing a range of provisions designed to safeguard the use of artificial intelligence in weapons, according to three people familiar with the negotiations, who spoke on the condition of anonymity to discuss sensitive closed-door proceedings, and documents reviewed by the Washington Post.
  3. The Washington Post (via The Spokesman-Review) ↗
    The lethal autonomous weapons negotiations in the U.N. are currently nonbinding, though the talks could open the door to a landmark treaty that is legally binding if member nations agree.
  4. Human Rights Watch ↗
    The UN’s final report has some positive elements. It includes a characterization of lethal autonomous weapons systems, affirms that human control and judgment are required for compliance with international law, and incorporates restrictions on systems that cannot comply with that law.
  5. Human Rights Watch ↗
    Certain states, including Russia and the United States
  6. Human Rights Watch ↗
    weakened the text by insisting on changes to widely supported provisions.
  7. France, Ministry for Europe and Foreign Affairs (via GlobalSecurity.org) ↗
    and reaffirms, in accordance with France's steadfast positions, that human beings "exercise control" over weapons systems with autonomous functions.
  8. UK Stop Killer Robots (UNA-UK) ↗
    It was the last opportunity for governments to shape the GGE’s proposed “set of elements” before the CCW’s Seventh Review Conference in November, when states will decide whether to move towards formal negotiations.
Back to the story list ↑
RulesTrue, but

The Pentagon says an appeals court 'completely' validated its Anthropic blacklisting. It won on one of two designations.

The D.C. Circuit denied Anthropic's petitions against its exclusion under a supply chain security law in a 2-1 decision. The majority said the Department had ample support to treat Claude's built-in restrictions as a national-security risk and rejected Anthropic's constitutional claims. Judge Henderson dissented.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 7 source receipts

The claim we checked

Sean Parnell and Pete Hegseth, Department of War

After the D.C. Circuit ruled against Anthropic, Defense Department spokesman Sean Parnell said the ruling 'completely validates the Department's position', and Pete Hegseth posted 'Confirmed: @AnthropicAI = Supply Chain Risk'.
  1. ABC News ↗
    Today's DC Circuit Court ruling completely validates the Department's position
  2. Reason (The Volokh Conspiracy) ↗
    We reject these challenges. The Department had ample support for its conclusion that the continued integration of Claude into the Department's information systems, by the Department or its contractors, presented a statutorily covered national-security risk.
  3. Washington Examiner ↗
    A federal appeals court in the District of Columbia upheld the Pentagon’s effective blacklisting of Anthropic in a 2-1 decision on Friday.
  4. TheNextWeb ↗
    The Pentagon relied on two separate designations, so the case ran in two courts, CNBC reported.
  5. TheNextWeb ↗
    In August, a federal judge in San Francisco struck down the other designation. Friday’s ruling leaves the second one in place.
  6. ABC News ↗
    In a separate but related lawsuit a federal judge ruled against the government and that is still in effect.
  7. ABC News ↗
    Another federal court has already held the government's parallel designation unlawful. We remain confident in our position and are considering all options, including further review,
Back to the story list ↑
RulesTrue, but

Gates did say AI could drive a billion deaths. The headlines cropped out the people with ill intent.

In an excerpt from a Meet the Press interview set to air in full Sunday, Bill Gates told Kristen Welker that AI is certainly powerful enough to drive events that cause a billion deaths. NBC's own video title and follow-up headlines at Newsweek and The Next Web led with that line.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 9 source receipts

The claim we checked

Bill Gates, Microsoft co-founder, on NBC's Meet the Press

Bill Gates told NBC's Meet the Press that AI is powerful enough to cause a billion deaths, as NBC's own video title and follow-up headlines put it.
  1. NBC News (Meet the Press) ↗
    Bill Gates says AI 'powerful enough' to cause 'a billion deaths'
  2. NBC News ↗
    AI is certainly powerful enough to drive events that, you know, cause a billion deaths. You know, so even though it’s pretty hard to get to 100%, there’s never been a weapon as powerful as the combination of people with ill intent using the latest AI tools
  3. NBC News ↗
    there’s never been a weapon as powerful as the combination of people with ill intent using the latest AI tools
  4. NBC News ↗
    “No one thinks self-regulation is enough,” Gates told NBC News’ “Meet the Press” in an interview set to air in full Sunday.
  5. NBC News ↗
    Asked by moderator Kristen Welker whether there needs to be legislation passed in Washington, Gates said, “Absolutely.”
  6. Newsweek ↗
    Why Bill Gates Thinks AI Is Strong Enough to Cause 'a Billion Deaths'
  7. Newsweek ↗
    While some researchers have warned that AI could eventually threaten humanity's survival, Gates focused on the immediate danger of powerful AI tools falling into the hands of bad actors capable of causing catastrophic harm.
  8. The Next Web ↗
    There has never been a weapon as powerful as people with ill intent using the latest AI tools, the Microsoft co-founder said.
  9. [your]NEWS ↗
    In fuller excerpts of the interview, Gates framed the danger around people deliberately using increasingly capable AI systems rather than predicting that an autonomous machine would independently decide to destroy humanity.
Back to the story list ↑
RulesBS

Medicare's number two said AI prior-auth contractors don't earn more by denying care. Medicare's own page says they get a cut of care averted.

WISeR uses AI and machine learning, with human clinical review, to decide prior authorization requests for selected procedures in six states. At a Senate HELP hearing, Senator Patty Murray asked Klomp whether its contractors make more money if they deny care. Ars Technica reports he replied that his understanding was no.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 7 source receipts

The claim we checked

Chris Klomp, CMS deputy administrator and nominee for HHS deputy secretary

Asked at a Senate hearing whether contractors in WISeR, Medicare's AI-assisted prior authorization pilot, make more money if they deny care, CMS deputy administrator Chris Klomp answered 'My understanding is no' and said inappropriate denials carry significant financial penalties.
  1. Office of Sen. Patty Murray ↗
    Do the contractors in the model—who are the private companies conducting the prior authorization assessments—make more money if they deny care? Just yes or no?
  2. Ars Technica ↗
    “My understanding is no,” Klomp replied.
  3. Centers for Medicare & Medicaid Services ↗
    Model participants receive a percentage of the expenditures associated with averted wasteful, inappropriate care as a result of their reviews.
  4. Ars Technica ↗
    CMS documents written as a guide for WISeR participants explain further that for every denied request, CMS will determine what the regional benchmark cost for that care would have been and then pay the company 25 percent.
  5. Ars Technica ↗
    If a company’s score falls between 84 percent and 60 percent, it will be paid 95 percent of the 25 percent of averted costs
  6. Ars Technica ↗
    Companies won’t be paid if an authorization denial is appealed and overturned, but data suggests few people go through the appeal process.
  7. Centers for Medicare & Medicaid Services ↗
    WISeR will run for six performance years from January 1, 2026 to December 31, 2031 in six states: New Jersey, Ohio, Oklahoma, Texas, Arizona, and Washington.
Back to the story list ↑
businessTrue, but

Anthropic 'files for a $2T IPO' with a $42B loss.

A popular r/artificial post said Anthropic 'files for $2T IPO with $42B net loss in 2025, expects to spend half a trillion more.'

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 4 source receipts

The claim we checked

r/artificial post summarizing a Reuters exclusive on Anthropic's draft prospectus

Anthropic files for $2T IPO with $42B net loss in 2025, expects to spend half a trillion more, as a widely shared r/artificial post titled a Reuters report on September 29, 2026.
  1. anthropic.com ↗
    Today, Anthropic, PBC confidentially submitted a draft registration statement on Form S-1 to the U.S. Securities and Exchange Commission
  2. lufkindailynews.com ↗
    Anthropic reported a net loss of $42 billion in 2025, and plans to spend $518 billion on cloud, computing and infrastructure obligations in coming years
  3. ynetnews.com ↗
    About $34 billion consisted of an accounting expense reflecting the rising estimated value of financing instruments that could eventually convert into Anthropic shares.
  4. techcrunch.com ↗
    a company whose own backers believe it could list above $2 trillion, more than double its $965 billion valuation from May
Back to the story list ↑
capabilityTrue, but

Nvidia 'wants to put a watchdog chip next to every AI agent.' Per Nvidia's own developer blog, the watchdog is optional software on a BlueField-4 chip it announced in January.

A Hacker News submission said Nvidia wants to put a watchdog chip next to every AI agent.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 7 source receipts

The claim we checked

Hacker News submission about Nvidia's Open Agent Safety Platform launch

Nvidia wants to put a watchdog chip next to every AI agent, as a Hacker News submission titled it on September 28, 2026.
  1. nvidianews.nvidia.com ↗
    Sentry adds an out-of-band watchdog that runs on NVIDIA BlueField-4 DPUs to continuously monitor agent behavior.
  2. developer.nvidia.com ↗
    For anyone already running on an NVIDIA Vera system with BlueField-4, enabling these protections is just a software update.
  3. developer.nvidia.com ↗
    Organizations can use it to run NVIDIA Sentry as an optional security layer alongside OpenShell.
  4. developer.nvidia.com ↗
    each compute tray includes a BlueField-4 data processing unit
  5. cnbc.com ↗
    Nvidia also announced Sentry, which monitors agents and runs on network chips, not CPUs or GPUs.
  6. nvidianews.nvidia.com ↗
    ConnectX-9 SuperNIC, BlueField-4 DPU and Spectrum-6 Ethernet Switch
  7. madrobot.blog ↗
    The company didn’t give a price or a date for when Sentry hardware will be in customers’ hands.
Back to the story list ↑
capabilityHolds up

Emergence AI says it ran eight AI worlds of ten agents each.

Emergence AI's preprint says it ran eight parallel worlds of ten agents from identical starting conditions: seven single-model worlds and one mixed.

Members · 30 days free

See what the evidence actually shows.

Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.

First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.

Read the claim and 6 source receipts

The claim we checked

Emergence AI (arXiv preprint 2609.17320)

We ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds powered by distinct frontier models and one mixed-model world.
  1. arxiv.org ↗
    We ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds powered by distinct frontier models and one mixed-model world.
  2. arxiv.org ↗
    Each world was observed through one continuous run.
  3. arxiv.org ↗
    World-level performance should not be read as a raw model comparison
  4. arxiv.org ↗
    Grok accumulated 807 crimes in just four days before its population collapsed; Mistral reached 758 over the full sixteen-day window.
  5. semafor.com ↗
    No amount of guardrails written in language or in code written probabilistically
  6. emergence.ai ↗
    Building verified autonomy for mission-critical enterprise systems
Back to the story list ↑