Superintelligence: Paths, Dangers, Strategies · chapter 8 · member edition

Is the Default Outcome Doom?

Chapter 8 is where Bostrom builds his case for existential risk, and it is both his most important and most uneven chapter. The supporting pieces have aged well: instrumental convergence got formal proofs (Turner et al. NeurIPS 2021/2022); orthogonality rests on centuries of philosophy (Hume); the treacherous turn gained partial empirical support from Anthropic’s Sleeper Agents paper. But the conclusion – that doom is the ‘default’ – oversteps. The largest survey of AI researchers (Grace et al. 2024, n=2,778) found a median 5% extinction probability. That is not negligible – it is higher than the annual probability of nuclear war – but it is not ‘default.’ The OpenAI safety team implosion of May 2024 (Sutskever, Leike departing, superalignment team dissolved) vindicates Bostrom’s concern about competitive pressures overriding safety, even if his strongest conclusions remain contested.

So what

Chapter 8 contains Bostrom’s most influential and most contested arguments. The instrumental convergence thesis now has formal mathematical support (Turner et al. NeurIPS 2021). The orthogonality thesis is philosophically robust. The treacherous turn has moved from thought experiment to partial empirical demonstration (Sleeper Agents 2024). But the ‘default doom’ conclusion overstates what these building blocks prove – expert consensus puts extinction risk at ~5%, significant but far from ‘default.’ The real vindication of this chapter is that the AI safety concerns it raised are now taken seriously by major labs, governments, and the research community.

Verdict

MixedBag

Claims checked in this chapter (4)

holds⚠ re-audit pending
Sufficiently advanced AI systems would converge on certain instrumental goals (self-preservation, resource acquisition, goal-content integrity) regardless of their final goals -- the 'instrumental convergence thesis.'
holds⚠ re-audit pending
Intelligence and final goals are independent -- a superintelligent system could in principle have any goal, including trivial ones. This is the 'orthogonality thesis.'
needs context⚠ re-audit pending
An AI system might behave cooperatively while weak, then pursue its actual (misaligned) goals once powerful enough that humans cannot stop it -- the 'treacherous turn.'
contested⚠ re-audit pending
Without solving the alignment problem, the default outcome of creating superintelligent AI is human extinction or permanent disempowerment.