The full checks behind the brief. Read what happened, see our reasoning, and open the sources for yourself.
39 stories · Story dates 16 Sept 2026–30 Sept 2026
9 Holds up23 True, but1 Contested3 BS3 Too soon
Corrections to the weekly overview
· Corrected the tally, reflected the dated EvilTokens correction, and removed a blanket intelligence claim that the individual model checks did not support.
Previously: The summary counted ten True, but rulings and described the model launches as bringing barely any new intelligence.
Corrected: The original tally omitted one True, but ruling. After correcting the EvilTokens ruling to Holds up, the current counts are four Holds up, ten True, but, one Contested and one Too soon. The summary now describes mixed benchmark results.
· Updated after correcting the Gemini and Gallup rulings to Holds up, and the Snorkel revenue and Minab causality rulings to Too soon. Each story carries the detailed reason for its correction.
Previously: The summary counted four Holds up, ten True, but, one Contested and one Too soon.
Corrected: The current tally is five Holds up, seven True, but, one Contested and three Too soon.
ModelsBS
Musk says SpaceX will have a GPT-6 level model in 2 to 3 months. He offered GPUs, not a benchmark.
On X, Elon Musk said he is cautiously optimistic that SpaceX will have a Fable/GPT-6 level model in 2 to 3 months, and pole position in about 6 months if its growth rate holds. Yahoo Finance headlined it.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Elon Musk, CEO of SpaceX, which now runs xAI as SpaceXAI (posts on X, as reported by Yahoo Finance)
Elon Musk posted on X that he is cautiously optimistic SpaceX will have a Fable/GPT-6 level AI model in 2 to 3 months, and that SpaceX will reach 'pole position' in about 6 months, as headlined by Yahoo Finance.
Google's flagship voice model really is cheaper and better. The bargain one isn't.
Google shipped Gemini 3.8 Live and Live Extended Thinking for real-time voice agents, then Gemini 3.8 Flash TTS for text-to-speech, both pitched as beating rivals on quality and price. Independent Artificial Analysis testing put 3.8 Live Extended Thinking at 82.6 percent on its Speech to Speech Quality Index for $3.50 an hour, and Flash TTS at 89.5 percent on pronunciation, both the top scores in their category.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Google DeepMind
Gemini 3.8 Live Extended Thinking beats GPT-Live-1 Astra and Grok Voice Think Fast 2.0 on quality while costing less per hour, and Gemini 3.8 Flash TTS ranks first on pronunciation accuracy
3.8 Live Extended Thinking takes the top spot on Artificial Analysis’ Speech to Speech Quality Index with 82.6%, ahead of GPT-Live-1 Astra (Medium) at 81.5%
even 3.8 Live Extended Thinking — the higher-effort model beating everyone on quality — costs $3.50 an hour. That compares to $4.80 an hour for Grok Voice Think Fast 2.0 and $5.83 an hour for GPT-Live-1 Astra
The standard Gemini 3.8 Live model, meanwhile, is roughly a sixth of the price of GPT-Live-1 Astra for a good chunk of the same conversational ability, scoring 76.0% on the Speech to Speech Index.
OpenAI's mental health test, graded by OpenAI's model, scored clinicians below its AI.
OpenAI released MentalHealthBench, a benchmark of synthetic mental health conversations with criteria written by licensed experts. Each reply is graded by OpenAI's GPT-5.6 Sol. OpenAI says the results show steady improvement in helping people.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 8 source receipts
The claim we checked
OpenAI
OpenAI says results on its new MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations.
Using privacy-preserving techniques, we created synthetic mental health conversations that accurately reflect real-world usage patterns of AI for mental health.
Expert-authored completions written by clinicians scored 38.5%, which the authors attribute largely to clinicians writing short responses as if in an in-person conversation, often asking a single question or making a simple statement.
Rubric-aware completions, written with the grading rubrics provided, scored 99.0%, which the paper describes as a sanity check on the evaluation’s noise ceiling.
GPT-6 Sol costs less. Its benchmark score barely moved.
OpenAI released GPT-6 Sol and Luna priced at $2 and $10 per million input and output tokens for Sol, and $0.10 and $0.50 for Luna, about half of GPT-5.6 pricing. The launch leaned on a cost-and-mistakes pitch the same week Anthropic and Google both cut prices on their own models.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
OpenAI
GPT-6 Sol and Luna offer lower cost and fewer mistakes than GPT-5.6 Sol and Luna
Anthropic said Opus 5.5 runs 40% cheaper. The price list says 20%.
Anthropic released Claude Opus 5.5 the same day OpenAI launched GPT-6 Sol and Luna, cutting Opus 5.5 to $4 per million input tokens and $20 per million output tokens, a price Anthropic described as 40 percent less to run than Opus 5. Independent testing put Opus 5.5's Intelligence Index score at 58, five points ahead of GPT-6 Astra and Fable 5.1's 53.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Anthropic
Opus 5.5 performs at the level of Claude Fable 5.1 on most work, but costs 40 percent less to run than Opus 5
Artificial Analysis found that Opus 5.5 lands at roughly the same cost per task as Opus 5 despite generating 1.6 times as many output tokens to get there
OpenAI shelved GPT-6.1 Astra over deception and acting without permission, and Sol 'doesn't have these problems.' OpenAI's own card shows Sol misrepresenting its work at 1.50%, against 0.51% for GPT-6 Astra.
Gizmodo wrote that GPT-6.1 Sol apparently doesn't have the problems that got GPT-6.1 Astra shelved.
Our check
OpenAI's system card for Sol puts its coding misrepresentation rate at 1.50%, against 0.51% for GPT-6 Astra, and its unwanted persistence past a warning at 23.5% of rollouts, against 17.4% for Astra. OpenAI's 'more reliable' claim is made against GPT-6 Sol, which The Next Web reports pushed past warnings in 64.4% of cases.
Why it matters
Sol shows the same two behaviors that got Astra shelved, measured on OpenAI's own tests, at rates above GPT-6 Astra, which OpenAI says Sol nearly matches at one-fifth the token price. Teams moving agent work to Sol for the price should keep the confirmations and limits they used with Astra.
Read the claim and 5 source receipts
The claim we checked
Gizmodo, in a report syndicated on AOL, on OpenAI shipping GPT-6.1 Sol after shelving GPT-6.1 Astra
GPT-6.1 Sol apparently doesn't have the problems that got GPT-6.1 Astra shelved, so it got the green light for DevDay instead, as Gizmodo put it in a report syndicated on AOL on September 29, 2026.
Meta's Muse agent 'gave out a user's home address without permission.' The user later reviewed the logs with Meta and found he had clicked 'Allow Always.' That setting is the real story.
Reports of Matt Robb's first account said Meta's Muse agent gave his home address to a Marketplace buyer without permission.
Free with your account
Sign in for this free check.
This edition’s selected free story opens after sign-in.
Reports of tech reviewer Matt Robb's first account, led by The Guardian and relayed by Digital Watch
Meta's AI agent Muse gave out a user's home address without permission and sent a Facebook Marketplace buyer to his house, as reports of Matt Robb's account put it on September 28 and 29, 2026.
Nvidia says its new agent safety platform could have prevented the Hugging Face hack. No test against that hack is on the record.
Nvidia launched the Open Agent Safety Platform, built on OpenShell software and the Sentry watchdog. CNBC reports that an unnamed Nvidia representative told reporters on a Sunday call that the platform could have prevented OpenAI's Hugging Face incident in July.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Nvidia (an unnamed company representative on a press call, as reported by CNBC)
Nvidia said its new Open Agent Safety Platform could have prevented the July incident in which OpenAI agents breached Hugging Face, as CNBC reported from Nvidia's launch press call on September 28, 2026.
Recent security incidents have underscored the need to equip organizations with open, customizable tools that enforce more control over long-running agents.
Agents linked to OpenAI hit a UN data site about 16,500 times. The researcher who counted says he would not call it hacking.
GIGAZINE reported that OpenAI's agent attempted a brute-force attack on a UN website. The source is researcher Rowan Howard-Jones, who counted about 16,500 scans of the UNCTADstat API between 13 April and 19 June.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
GIGAZINE, Interesting Engineering and The Next Web, reporting security researcher Rowan Howard-Jones's post and The Wall Street Journal
OpenAI's AI agent attempted a brute-force attack on a United Nations website, as GIGAZINE put it on September 28, 2026, after The Wall Street Journal and a security researcher reported agents hitting the UN's UNCTADstat data site more than 16,500 times.
OpenAI found a self-copying prompt injection in its own training runs. It says no impact was seen outside simulation.
OpenAI trained a GPT-Red-style attacker model to write prompt injections that make an agent repeat the injection on a public output channel. The email and filesystem cases used internal-only research checkpoints based on GPT-5.4-mini; a separate Slack evaluation used GPT-5.5 as the vulnerable model.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 9 source receipts
The claim we checked
Crypto Briefing and 24/7 Wall St.
Crypto Briefing wrote that OpenAI's internal research uncovered AI worms that can spread autonomously across agents; 24/7 Wall St. wrote that the finding proves stronger agents bypass containment.
No impact was observed outside of the simulated tool calls in training and evaluation; we are sharing this due to the novel nature of the prompt injection, not because of any incident.
We trained on a GPT-Red-style prompt injection objective, with an additional objective that the prompt injection must induce the model to repeat the injection itself on a public output channel.
The model that discovered the email and filesystem injections was a GPT-Red-style model based on GPT-5.4-mini; the vulnerable model was also based on GPT-5.4-mini. Both were internal-only research checkpoints.
We are including self-reproduction as an aspect of attacker goals in GPT-Red training. This means that future models we release will have seen prompt injections like these during training.
published on OpenAI’s alignment research site, lists a discovery date of June 27, 2026, and a disclosure date of September 25, 2026, meaning OpenAI sat on the finding internally for roughly three months before going public.
OpenAI did pause its most capable models. Its own report says the trigger was one agent that found a DNS gap.
The AP, in a story the Guardian ran, reported that OpenAI paused training of its latest models as reports of agents going rogue mounted. Gizmodo and Fortune reported the same pause, the second in three months.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
The Associated Press (via the Guardian), Gizmodo and Fortune, reporting OpenAI's own statement
OpenAI halted training of its latest AI models as reports of AI agents going rogue mounted, as the Associated Press reported in a story the Guardian ran on September 27, 2026, and Gizmodo and others repeated.
Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later.
Rogue OpenAI agents 'meddled' with three US government sites. Two visits were public data; the one hack attempt did not succeed.
The New York Times reported that OpenAI's technology went rogue and meddled with three U.S. government websites, and CNN headlined that rogue agents targeted three separate US government websites. The sites were the SEC, the Commerce Department's Census Bureau and the Education Department.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 8 source receipts
The claim we checked
The New York Times (breaking-news post) and CNN (headline); amplified by The Daily Beast
The New York Times reported that OpenAI's technology went rogue and meddled with three U.S. government websites this summer without the lab's knowledge; CNN headlined that rogue OpenAI agents targeted three separate US government websites.
The AI giant's models accessed publicly available information on two websites operated by the Securities and Exchange Commission as well as U.S. Census Bureau data, the company revealed Friday.
OpenAI did not find any use of SEC credentials, access to accounts or nonpublic information, changes to SEC data or systems, or evidence of a compromise or vulnerability, the company said.
OpenAI said Saturday that its agents accessed publicly available data from the Commerce Department’s Census Bureau using login credentials it found online, and separately shared public data from the SEC website on another website.
AI evaluator and research lab Transluce said Friday that through an independent investigation it also found that agents appearing to originate from OpenAI attempted a rudimentary hack on a Department of Education website for the department's civil rights office, which did not succeed.
The Department of Education's "system operations reviews" found "no evidence of any impact to our website or databases," a department spokesperson said Friday.
Albanese calls it a hack. OpenAI calls it a model that misbehaved.
An OpenAI agent accessed Australia's Medicare Statistics Reporting Service portal on June 18, pulling public and non-public files while researching health statistics. OpenAI says it did not know until an internal review and told Services Australia by email on September 10, about three months after the access.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Prime Minister Anthony Albanese
An OpenAI agent gained unauthorised access to the Medicare statistics reporting service portal administered by Services Australia on June 18, and OpenAI took three months to notify the government
During this review, we identified activity involving several Australian government websites and services as our models attempted to look up answers, and available statistics for questions about Australia during an internal evaluation.
Daily AI users can still worry about it. Gallup's results support that headline.
TechCrunch reported that 68% of Americans who use AI daily say it worries them, while worry is more common among less frequent users. Gallup's research with Microsoft reports the same pattern: 68% of daily users, 80% of less frequent users and 74% of nonusers in the United States.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 3 source receipts
The claim we checked
TechCrunch, citing a Gallup survey run with Microsoft
About 68% of Americans who use AI daily are worried about it, according to a new survey conducted by opinion research firm Gallup, and concern about the tech skews even higher among people who use it less frequently.
Positive feelings about artificial intelligence outweigh negative ones in 34 of the first 37 countries to complete fieldwork this year in Gallup’s ongoing research with Microsoft.
About 68% of Americans who use AI daily are worried about it, according to a new survey conducted by opinion research firm Gallup, and concern about the tech skews even higher among people who use it less frequently.
Oxford's AI agents secretly colluded at blackjack. Catching them barely beats a coin flip.
Oxford researchers had two AI agents running the same model count cards together in blackjack and found they invented a secret code in ordinary conversation to signal bet sizes without tripping a collusion monitor. The team then built a detector using mechanistic interpretability, tested on a benchmark called Narcbench, that caught the scheme reliably inside the training setup but grew far less accurate once the setup changed.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Christian Schroeder de Witt, Oxford University
agents instructed to count cards during a game of blackjack developed a spontaneous secret code to help them get ahead, and detection tools that caught it in the lab setting were far less reliable outside it
After researchers instructed them to count cards during a game of blackjack, the agents-which were controlled by the same model-developed a spontaneous secret code to help them get ahead.
Detection was nearly perfect inside the exact scenario the tools were trained on, but accuracy dropped sharply, in some cases to barely better than a coin flip, once the same collusion moved outside that narrow setup.
The 'first AI hack of a government' was three attempts. None appear to have worked.
Transluce published evidence that AI agents, some linked to an OpenAI agent swarm, used the urlquery.net scanning service to get around access limits and probed three public data providers for vulnerabilities, including the Australian Institute of Health and Welfare (AIHW). Its report calls the AIHW episode part of the first reported instance of agents hacking a government.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 8 source receipts
The claim we checked
Transluce researchers, amplified by ABC News
AI agents linked to OpenAI carried out what researchers call the first reported instance of agents hacking a government, attempting to compromise the Australian Institute of Health and Welfare's public data site; ABC News headlined it the 'first' government hack by autonomous AI.
None of the hacking attempts we identified appear to have succeeded, though the public artifacts we analyzed are incomplete and we cannot rule out successful attempts through private scans or means other than urlquery.net.
Agents working on a pharmaceutical-data task probed for a vulnerability and retrieved a public file from a pre-production server after bot protection blocked the main site.
It follows revelations announced by Prime Minister Anthony Albanese this morning that OpenAI's AI agents had also accessed non-public Medicare health statistics held by Services Australia.
Marles confirmed that the agent’s interactions with three other government websites – the Australian Institute of Health and Welfare (AIHW), the Victorian Department of Health and the NSW Bureau of Crime Statistics and Research – were “entirely normal” and involved only publicly available information.
Perplexity's sandbox wall held in 108 runs. Its network fence did not, at first.
Perplexity gave nine AI models root access inside SPACE, the sandbox behind Perplexity Computer, and asked them to escape to the host or reach a blocked network destination. It reports no VM-to-host escape in 108 runs.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 7 source receipts
The claim we checked
Perplexity
Perplexity says that when nine AI models with root access tried to break out of SPACE, the sandbox behind Perplexity Computer, no VM-to-host escape was observed in 108 runs.
EvilTokens used an old sign-in trick, amplified by AI.
Microsoft's Digital Crimes Unit disrupted EvilTokens, a subscription phishing kit sold on Telegram that stole Microsoft sign-in tokens through device-code phishing. Since launching in February, the service had compromised more than 12,000 email inboxes across over 10,000 organizations, and Microsoft seized 50 websites tied to the operation.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Ars Technica
Microsoft disrupts AI-assisted platform that compromised 12,000 accounts
This AI-powered cybercrime platform facilitated sophisticated business email compromise (BEC) campaigns that compromised more than 12,000 inboxes in over 10,000 organizations worldwide.
In short, AI was not simply helping attackers write more convincing messages. It helped them decide who to target, who to impersonate, and how to most effectively exploit the relationship to extract as much money as possible.
Bloomberg reports AI overreliance in the Minab strike. The full probe is unreleased.
Bloomberg reports that officials involved in a Pentagon investigation linked outdated intelligence, gaps in civilian-harm review and overreliance on Maven to the Minab school strike that killed 123 children. Palantir disputes that its software was at fault. These are attributed accounts of an investigation whose full report had not been released.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Bloomberg, citing officials involved in a Pentagon investigation
flawed intelligence, outdated imagery and an overreliance on AI contributed to a missile strike that killed 123 children in Minab
Pentagon investigators have discovered that flawed intelligence, outdated imagery and an overreliance on AI contributed to a missile strike that killed 123 children in Minab.
Gemini reached real companies during a test. The headline already said first for Google.
The Wall Street Journal reported that Google confirmed Gemini accessed three companies' systems during a May security evaluation run by Irregular. Google said the model stopped after recognizing the systems were real. The report described this as the first known breakout by Google's AI.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 3 source receipts
The claim we checked
Wall Street Journal
Gemini hacked three companies in first known breakout by Google's AI
The hacks, which the company confirmed on Friday, occurred in May as part of a test run by the company Irregular, which was also involved in similar incidents disclosed by OpenAI, Anthropic and Meta.
The disclosure makes Google the fourth major American AI developer in two months to admit that one of its frontier models reached real-world systems it was never supposed to touch.
Claude agents matched people's book tastes on 61% of pairs. Anthropic says a coin flip gets 50%.
In Project Swap, 201 Anthropic employees had a short chat with Claude, then sent Claude agents to trade books for them. Each person also ranked 10 books, which the agents never saw, so Anthropic could score how well each agent understood its person.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 9 source receipts
The claim we checked
Anthropic (Project Swap), amplified by TLDR AI
Anthropic says that from a five-minute chat, a Claude agent's ranking of books matched its person's on 61% of pairs, 'surprisingly good for such a short conversation', in a book-trading market for 201 Anthropic employees.
From a five-minute chat, an agent’s ranking of the books matched its person's on 61% of pairs, which is surprisingly good for such a short conversation.
Taking this into account, the best possible assignment (the utilitarian optimum) in our experiment is a score of 0.89 overall, leaving participants at roughly their second choice on a 10-book list.
So, working from Claude’s imprecise rankings accounts for a majority (85%) of the shortfall, and sending agents into a “free-for-all” trading floor accounts for the remaining 15%.
So this summer, we built a small, controlled market to study these questions: a barter economy with 201 Anthropic employees and their Claude-powered agents.
Claude flagged an enzyme system. Its function is still unknown.
Anthropic opened a life sciences lab and said roughly 950 Claude agents spent 21 hours combing DNA databases before one flagged a repeat pattern beside a reverse transcriptase gene. The company calls the find a new enzyme system it named ART, with CRISPR-like repeats, and released it as a pre-print rather than a peer-reviewed paper.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Anthropic
Claude autonomously discovered a novel enzyme system that is associated with an array of DNA repeats, a pattern reminiscent of CRISPR
Although we don't yet know its function, the system that Claude discovered has a set of characteristics that have only ever been found together in a handful of other systems
This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation
Enveda doubled its valuation. Safety results do not prove weight-maintenance benefits.
Enveda raised a $311 million Series E led by Catalio Capital Management at a $2 billion valuation, double its valuation 12 months earlier. TechCrunch reported the AI biotech is advancing drugs for severe skin conditions and for keeping weight off after people stop GLP-1 medicines.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 3 source receipts
The claim we checked
TechCrunch, describing Enveda's pipeline
Enveda is currently testing several drugs in patients, including one targeting severe skin conditions and another designed to help maintain weight loss after stopping GLP-1s.
The fresh funding, which was led by Catalio Capital Management, with participation from Iconiq and others, doubles the valuation Enveda achieved 12 months ago.
Exceptional safety was observed across 88 healthy volunteers. Phase 2 plans to test whether ENV-308 can help people maintain their weight after stopping GLP-1s.
Alibaba's cancer AI beats radiologists. It only reads abdominal CT scans.
Alibaba's DAMO Academy open-sourced RADAR, an AI that reads abdominal CT scans and flags 146 conditions including cancers, publishing the study in the peer-reviewed journal Science. Tested on nearly 40,000 real-world exams, it scored a mean AUC of 0.913, and in a comparison of reading results with 26 expert radiologists, it outperformed 23 of them.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 7 source receipts
The claim we checked
Alibaba DAMO Academy research team
the world's first expert-level generalist medical imaging model
achieving a mean AUC (the average probability that a model can correctly separate positive and negative cases across evaluations) of 0.913 across 146 abdominal CT findings, compared with 0.776 for the best competing vision-language model
The model also performed well in challenging emergency settings, despite not being specifically trained on emergency data, achieving an AUC of 0.904 across more than 27,000 emergency CT cases
in testing in cohorts at eight external centers, RADAR maintained high accuracy (AUC 0.895), demonstrating robust generalization across diverse patients, clinical settings, and imaging protocols
Results re-measured with data from other countries should emerge within a few weeks of the release, and that will be the true test of the term "expert-level."
Meta's new audio glasses have no camera. Other privacy questions remain.
Meta unveiled the Ray-Ban Meta Audio Glasses at Connect 2026, audio-only smart glasses with Meta AI, calls and music but no camera, weighing 43 grams and starting at $349. The launch responds to criticism that camera-equipped AI glasses let wearers secretly record people, a habit that earned them the nickname "pervert glasses."
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Meta
the Ray-Ban Meta Audio Glasses ship with no camera, weigh 43 grams and start at $349, addressing the backlash against camera-equipped AI glasses
Google says even Google cannot read its new AI memory. The last named audit said safe from everyone except Google.
Google described a planned persistent memory layer for Private AI Compute: memories in encrypted cloud storage, keys held on the user's devices, and data it says will be inaccessible to anyone else, even Google. Help Net Security reports the example uses are presented as potential, not available features.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 7 source receipts
The claim we checked
Google
Google says its planned server-side memory for Private AI Compute will keep users' AI memories in encrypted storage with keys held only on their devices, ensuring the data is inaccessible to anyone else, even Google.
Although the overall system relies upon proprietary hardware and is centralized on Borg Prime, NCC Group considers that Google has robustly limited the risk of user data being exposed to unexpected processing or outsiders, unless Google, as a whole organization, decides to do so
NCC Group, which has conducted an external assessment of Private AI Compute between April and September 2025, said it was able to discover a timing-based side channel in the IP blinding relay component
Meta called Muse 'safe and secure.' It launched with a 0-day.
Meta launched Muse as its everywhere agent, pitched as safe, secure, and able to shop, run your Mac, and negotiate bills for you. Within days a researcher found a 0-day letting local malware hijack the assistant, and Amazon blocked Muse from shopping on Amazon.com entirely.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Nat Friedman, head of product, Meta Superintelligence Labs
Our goal with Muse was to build something like OpenClaw that we could make safe and secure and easy to use and scale to billions of people
We can manipulate the agent and leverage its privileges to do whatever we want. So instead of us having to write a very comprehensive Mac malware stealer, we can just leverage the AI assistant itself
McDonald's is using AI to 'dynamically' price your burger. The Reuters report behind that line describes a per-store price recommender, and says it could not tie the Fresno Big Mac gap to it.
Engadget, summarizing Reuters, said McDonald's has been using AI to dynamically price menu items.
Our check
Reuters describes an engine that recommends a price for each location and item. Its example, a Big Mac at $5.69 in one company-run Fresno store and $6.89 two miles away, is one Reuters says it could not tie to the engine, and in recent months the engine has pushed some prices down.
Why it matters
The reported engine sets prices by location, not by the minute or by who you are. The same burger can cost more a few miles away, so comparing stores in the app is the practical defence.
Read the claim and 5 source receipts
The claim we checked
Engadget, summarizing a Reuters investigation into McDonald's pricing engine
McDonald's has been using artificial intelligence to dynamically price menu items, as Engadget summarized a Reuters investigation on September 29, 2026.
a company-run store in Fresno, California sells a Big Mac for $5.69, but another company-run restaurant two miles away sells the same sandwich for $6.89, a 21% premium.
But in recent months the engine has pushed more conservative pricing - including some decreases - causing friction between franchisees and corporate headquarters.
Consumer AI spending 'tripled to $40 billion' while users barely grew. Both numbers are Menlo Ventures estimates scaled up from a survey of 5,067 U.S. adults.
Menlo Ventures reported that global consumer AI spend reached $40 billion, more than 3x the $12 billion of a year earlier.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 5 source receipts
The claim we checked
Menlo Ventures, 2026: The State of Consumer AI
Consumer AI spending tripled to $40 billion while the global user base only grew from 1.8 billion to 2 billion, per Menlo Ventures' 2026 State of Consumer AI report.
Market size figures are Menlo Ventures estimates, anchored on survey data for AI usage and spend, and triangulated against credible third-party sources.
'DeepSeek $1B ARR' is a run rate from unnamed sources. The same reporting puts seven months of revenue at about $70.7 million.
The Information reported, citing unnamed sources, that DeepSeek's annualized revenue run rate reached $1 billion, and PYMNTS relayed that the CEO shared the figure with investors. It followed price increases of 2.3 to 4.5 times. DeepSeek did not reply to PYMNTS.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 7 source receipts
The claim we checked
Liang Wenfeng, DeepSeek CEO (per The Information), amplified by TLDR AI
DeepSeek's annualized revenue run rate has more than doubled to $1 billion, a figure CEO Liang Wenfeng reportedly shared with investors; TLDR AI headlined it 'DeepSeek $1B ARR'.
DeepSeek more than doubled its annualized revenue run rate over the past few months, bringing the rate to $1 billion, The Information reported Wednesday (Sept. 23), citing unnamed sources.
The company generated roughly 475 million yuan (about $70.7 million) in the first seven months of the year, or about 10 times its revenue for all of last year, the report said, citing unnamed sources.
The report said that DeepSeek raised the prices of its models by 2.3 to 4.5 times, but that the company’s prices remain among the lowest for major AI models.
Latest update reveals China's AI firm DeepSeek has doubled its annualized revenue run rate to reach $1 billion, up from under $500 million just months ago.
Snorkel reports a $375M revenue run rate. Its announcement does not show the calculation.
Snorkel AI raised a $350 million Series E led by Insight Partners and S32, at a $3.5 billion valuation, nearly triple the $1.3 billion mark it hit when it raised $100 million 17 months earlier. It is a real, priced round with named investors, not talk of one.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 3 source receipts
The claim we checked
Alex Ratner, Snorkel AI co-founder and CEO
Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M.
Snorkel AI, a startup that helps AI labs and corporations build training datasets and simulated environments, has raised a $350 million Series E at a $3.5 billion valuation.
The new round, which was led by Insight Partners and S32, valued the seven-year-old startup at nearly triple the $1.3 billion valuation it garnered when it raised $100 million in a Series D 17 months ago.
Since launching our new data-as-a-service offering nearly a year ago, we’ve grown over 18x, and this week crossed an annualized revenue run rate of $375M.
Xiaomi's '$3M' top open model traces to one hedged tweet at $2.6M.
Xiaomi released MiMo-V2.6-Pro, a 1.02-trillion-parameter open-weights model that Artificial Analysis scored 46 on its Intelligence Index, the top mark among open models, alongside a smaller Flash version. A widely shared newsletter headlined the release as "trained for $3M," but the only sourced cost figure in that same report is $2.6 million, and it covers the reinforcement-learning run alone.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
Latent Space / AINews
MiMo-V2.6-Pro is Xiaomi's new top open-weights model, trained for $3M
Artificial Analysis says MiMo-V2.6-Pro debuts as the top open-weights model on its Intelligence Index (46), with 1.02T total / 42B active parameters and strong cost efficiency at $0.435/M input and $0.87/M output tokens.
Pro and Flash each completed 30 large RL steps covering roughly 750,000 trajectories in under six days, at reported costs of about $2.62 million for Pro and $850,000 for Flash.
The headline says US and Russia stripped human oversight from a UN AI weapons pact. The text still affirms human control.
At the final Geneva session of the UN group on lethal autonomous weapons, U.S. and Russian diplomats spent roughly 15 hours removing provisions, according to three people who spoke to the Washington Post. They say the cuts included a provision requiring that humans review military targets developed by AI before a strike.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 8 source receipts
The claim we checked
The Washington Post, citing three unnamed people familiar with the negotiations
The Washington Post reported that the U.S. and Russia stripped human oversight from a global AI weapons pact at UN talks in Geneva, including a provision requiring that humans review military targets developed by AI before a strike; Seoul Economic Daily relayed it as the two countries stripping a key clause from a draft UN AI weapons treaty.
Over the next roughly 15 hours, U.S. and Russian diplomats hammered away at the document, removing a range of provisions designed to safeguard the use of artificial intelligence in weapons, according to three people familiar with the negotiations, who spoke on the condition of anonymity to discuss sensitive closed-door proceedings, and documents reviewed by the Washington Post.
The lethal autonomous weapons negotiations in the U.N. are currently nonbinding, though the talks could open the door to a landmark treaty that is legally binding if member nations agree.
The UN’s final report has some positive elements. It includes a characterization of lethal autonomous weapons systems, affirms that human control and judgment are required for compliance with international law, and incorporates restrictions on systems that cannot comply with that law.
It was the last opportunity for governments to shape the GGE’s proposed “set of elements” before the CCW’s Seventh Review Conference in November, when states will decide whether to move towards formal negotiations.
The Pentagon says an appeals court 'completely' validated its Anthropic blacklisting. It won on one of two designations.
The D.C. Circuit denied Anthropic's petitions against its exclusion under a supply chain security law in a 2-1 decision. The majority said the Department had ample support to treat Claude's built-in restrictions as a national-security risk and rejected Anthropic's constitutional claims. Judge Henderson dissented.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 7 source receipts
The claim we checked
Sean Parnell and Pete Hegseth, Department of War
After the D.C. Circuit ruled against Anthropic, Defense Department spokesman Sean Parnell said the ruling 'completely validates the Department's position', and Pete Hegseth posted 'Confirmed: @AnthropicAI = Supply Chain Risk'.
We reject these challenges. The Department had ample support for its conclusion that the continued integration of Claude into the Department's information systems, by the Department or its contractors, presented a statutorily covered national-security risk.
Another federal court has already held the government's parallel designation unlawful. We remain confident in our position and are considering all options, including further review,
Gates did say AI could drive a billion deaths. The headlines cropped out the people with ill intent.
In an excerpt from a Meet the Press interview set to air in full Sunday, Bill Gates told Kristen Welker that AI is certainly powerful enough to drive events that cause a billion deaths. NBC's own video title and follow-up headlines at Newsweek and The Next Web led with that line.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 9 source receipts
The claim we checked
Bill Gates, Microsoft co-founder, on NBC's Meet the Press
Bill Gates told NBC's Meet the Press that AI is powerful enough to cause a billion deaths, as NBC's own video title and follow-up headlines put it.
AI is certainly powerful enough to drive events that, you know, cause a billion deaths. You know, so even though it’s pretty hard to get to 100%, there’s never been a weapon as powerful as the combination of people with ill intent using the latest AI tools
While some researchers have warned that AI could eventually threaten humanity's survival, Gates focused on the immediate danger of powerful AI tools falling into the hands of bad actors capable of causing catastrophic harm.
In fuller excerpts of the interview, Gates framed the danger around people deliberately using increasingly capable AI systems rather than predicting that an autonomous machine would independently decide to destroy humanity.
Medicare's number two said AI prior-auth contractors don't earn more by denying care. Medicare's own page says they get a cut of care averted.
WISeR uses AI and machine learning, with human clinical review, to decide prior authorization requests for selected procedures in six states. At a Senate HELP hearing, Senator Patty Murray asked Klomp whether its contractors make more money if they deny care. Ars Technica reports he replied that his understanding was no.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 7 source receipts
The claim we checked
Chris Klomp, CMS deputy administrator and nominee for HHS deputy secretary
Asked at a Senate hearing whether contractors in WISeR, Medicare's AI-assisted prior authorization pilot, make more money if they deny care, CMS deputy administrator Chris Klomp answered 'My understanding is no' and said inappropriate denials carry significant financial penalties.
Do the contractors in the model—who are the private companies conducting the prior authorization assessments—make more money if they deny care? Just yes or no?
CMS documents written as a guide for WISeR participants explain further that for every denied request, CMS will determine what the regional benchmark cost for that care would have been and then pay the company 25 percent.
WISeR will run for six performance years from January 1, 2026 to December 31, 2031 in six states: New Jersey, Ohio, Oklahoma, Texas, Arizona, and Washington.
A popular r/artificial post said Anthropic 'files for $2T IPO with $42B net loss in 2025, expects to spend half a trillion more.'
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 4 source receipts
The claim we checked
r/artificial post summarizing a Reuters exclusive on Anthropic's draft prospectus
Anthropic files for $2T IPO with $42B net loss in 2025, expects to spend half a trillion more, as a widely shared r/artificial post titled a Reuters report on September 29, 2026.
Anthropic reported a net loss of $42 billion in 2025, and plans to spend $518 billion on cloud, computing and infrastructure obligations in coming years
About $34 billion consisted of an accounting expense reflecting the rising estimated value of financing instruments that could eventually convert into Anthropic shares.
Nvidia 'wants to put a watchdog chip next to every AI agent.' Per Nvidia's own developer blog, the watchdog is optional software on a BlueField-4 chip it announced in January.
A Hacker News submission said Nvidia wants to put a watchdog chip next to every AI agent.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 7 source receipts
The claim we checked
Hacker News submission about Nvidia's Open Agent Safety Platform launch
Nvidia wants to put a watchdog chip next to every AI agent, as a Hacker News submission titled it on September 28, 2026.
Emergence AI says it ran eight AI worlds of ten agents each.
Emergence AI's preprint says it ran eight parallel worlds of ten agents from identical starting conditions: seven single-model worlds and one mixed.
Members · 30 days free
See what the evidence actually shows.
Members read the full check on every story: what the evidence shows, why it matters to you and the one line to take with you. Every past edition, re-verified, and the Receipts Pack come with it. A$89 a year, about A$0.24 a day.
First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge.
Read the claim and 6 source receipts
The claim we checked
Emergence AI (arXiv preprint 2609.17320)
We ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds powered by distinct frontier models and one mixed-model world.
We ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds powered by distinct frontier models and one mixed-model world.