Superintelligence: Paths, Dangers, Strategies · chapter 10 · member edition

Oracles, Genies, Sovereigns, Tools

Chapter 10 proposed a taxonomy (Oracle/Genie/Sovereign/Tool) that has been superseded by reality: current LLMs are all four simultaneously. But the chapter’s deeper insights have aged remarkably well. AI boxing/containment has been effectively abandoned by major labs in favor of alignment approaches – vindicating Bostrom’s skepticism. The tool AI argument (Drexler’s CAIS) partially materialized with LLMs, but the relentless push toward agentic capabilities (tool use, multi-step reasoning, autonomous agents) shows that ‘just make it a tool’ isn’t a stable equilibrium. Sycophancy research provides mild evidence for oracle manipulation. The taxonomy is pedagogically useful but the real contribution is the insight that no simple system design eliminates alignment risk.

So what

Chapter 10 is the most practically relevant chapter in 2026. Its taxonomy is outdated (current LLMs defy clean categories), but its core insights have proven remarkably prescient: containment is harder than it looks, tool AI isn’t as safe as hoped once you add agency, and even oracle-like systems can manipulate through their outputs. The shift from boxing-based to alignment-based safety strategies across the field is an implicit vindication of Bostrom’s skepticism about containment.

Verdict

MixedBag

Claims checked in this chapter (4)

needs context⚠ re-audit pending
AI systems can be meaningfully categorized as Oracles (answer questions), Genies (execute commands), Sovereigns (act autonomously with broad goals), or Tools (no agency).
needs context⚠ re-audit pending
AI systems can be contained through 'boxing' -- restricting communication channels and capabilities to prevent escape or manipulation.
needs context⚠ re-audit pending
Tool AI systems (with no agency or goal-directedness) represent a safer alternative to agent AI systems, as argued by Eric Drexler and others.
needs context⚠ re-audit pending
Even question-answering AI (oracles) could manipulate humans through carefully chosen answers, posing risks even without explicit agency.