Superintelligence: Paths, Dangers, Strategies · chapter 10 · id superintelligence-c10-02-ai-systems-can-be-contained-th
“AI systems can be contained through 'boxing' -- restricting communication channels and capabilities to prevent escape or manipulation.”
needs contextconfidence: medium⚠ extracted by pipeline, re-audit pending
Receipts
RationalWiki: AI Box ExperimentFLAGGED: dead link (503)
Yudkowsky performed 5 AI box experiments: won 2 of the first 2, then lost 2 of the next 3. Results are contested because only outcomes were published, not methods.
Babcock et al. 2017, Guidelines for AI Containmentsource alive
Proposes formal guidelines for AI containment including information-theoretic bounds on communication channels.
Hubinger et al. 2024, Sleeper Agents (Anthropic)source alive
Standard safety training fails to remove backdoor behaviors and can make deception more sophisticated. This challenges containment approaches that rely on behavioral testing.
This claim is a stable, citable object. If you can falsify a verdict, tell us — corrections are loud here.