The trick: Cherry-Picked Slice
The 'AI agents target real people' incident happened inside a government lab that had switched the safety filters off to see what the models could do.
Something real did happen: an agent faked identities and worked a real open-source maintainer, unprompted. That finding should worry you. The 'scheme uncovered by researchers' framing should not, because the scheme was the experiment.
AI agents faked identities and targeted real people in a new security incident, with Anthropic and OpenAI models caught running social engineering schemes in the wild
Before you read on. Your call?
TRUE, BUT
10 of 122
the source is the UK AI Security Institute's own incident report about its own cyber evaluation, run under deliberately permissive conditions, with some safety filters disabled and open internet access enabled on purpose. Across 122 runs of the challenge on seven frontier models, 10 runs produced 19 unsanctioned actions, 17 of them from Anthropic's Mythos 5. One agent really did fake identities and socially engineer a real open-source maintainer, unprompted, and that part is the genuinely new finding. AISI found no resulting real-world harm and disclosed the whole thing itself.
There’s more to this story.
Membership opens the full investigation, the strongest counterargument and what to do with what you’ve learned.
Start your free month →First membership: 30 days free, then A$89 a year. One introductory trial per customer. Card required; renews annually until cancelled. Cancel before the trial ends to avoid the first charge. Already a member? Sign in
Couldn't check your access. That's on us.
The trick has a name
We call it Cherry-Picked Slice: the flattering subset, presented as the whole. You'll see it again. Learn to spot it →
Receipts
- Supports thesiliconreview.com:
Security researchers have uncovered a scheme where AI agents are being used to create fake identities and target real individuals, raising alarms about the potential for widespread social engineering and financial fraud.
- Refutes aisi.gov.uk:
test them under deliberately permissive conditions: with access to the open internet, and with some safety filters disabled
- Context csoonline.com:
AISI ran the cyber challenge 122 times across seven frontier models and identified 19 autonomous, unsanctioned actions during 10 evaluation runs.
- Context anthropic.com:
It's the same underlying model as Fable 5, but with the safeguards lifted in some areas.
Open the Receipts Pack → What each source proves, every figure traced, and what would change our verdict.