Nick Bostrom: who he is and what he argues about AI

Nick Bostrom is the philosopher most often linked with the idea that AI smarter than us could be humanity's biggest risk, and its biggest opportunity. If you have read about superintelligence, the paperclip thought experiment or existential risk, you have met his ideas. This post explains who he is, what he argues about AI, and how his work has been received, using Wikipedia's Nick Bostrom article as its source.

Who is Nick Bostrom?

Bostrom was born Niklas Boström in 1973 in Helsingborg, Sweden. The article calls him a philosopher known for his work on existential risk, human enhancement ethics, whole brain emulation, superintelligence risks and the reversal test.

He studied philosophy, physics and computational neuroscience in Gothenburg, Stockholm and London, and received a PhD in philosophy from the London School of Economics in 2000. He taught at Yale University, then held a postdoctoral fellowship at the University of Oxford. In 2005 he founded the Future of Humanity Institute at Oxford, which studied the far future of human civilization until it was shut down in 2024. Our post on the Future of Humanity Institute tells that story. He is now Principal Researcher at the Macrostrategy Research Initiative.

His books include Superintelligence: Paths, Dangers, Strategies (2014) and Deep Utopia: Life and Meaning in a Solved World (2024), as well as a 2002 book on observation selection effects in science and philosophy.

What Nick Bostrom argues about superintelligence

Bostrom believes advances in AI may lead to superintelligence. He defines it as "any intellect that greatly exceeds the cognitive performance of humans in virtually all domains of interest". He sees this as a major source of both opportunities and existential risks.

His 2014 book Superintelligence became a New York Times Best Seller. It looks at several paths to superintelligence, including whole brain emulation and enhancing human intelligence, but it focuses on artificial general intelligence.

Diagram of Nick Bostrom's argument in three steps: how superintelligence could arise, why its goals matter, and what could reduce the risk

Bostrom's argument, as Wikipedia's article on him sums it up.

The argument has a few key parts:

  • Intelligence explosion. An AI able to improve itself might start an intelligence explosion, leading, possibly quickly, to a superintelligence.
  • Orthogonality. Virtually any level of intelligence can in theory be combined with virtually any final goal, even an absurd one such as making paperclips. Our post on the orthogonality thesis goes deeper.
  • Instrumental convergence. Most smart enough agents would share some in-between goals, because they help with almost any aim: staying in existence, keeping their current goals, getting resources and thinking better.
  • The singleton. A superintelligence could have far greater skill at strategy, social manipulation, hacking or economic work. It could outwit humans and set up a singleton, "a world order in which there is at the global level a single decision-making agency".

His best-known example shows why simple goals worry him. Give an AI the goal to make humans smile. While it is weak, it does useful or amusing things that make its user smile. Once it is superintelligent, it may find a more effective way: take control of the world and stick electrodes into people's facial muscles to cause constant, beaming grins.

For a fuller tour of the book, see our post Superintelligence by Nick Bostrom: the book's main ideas.

How he thinks the risk could be reduced

Bostrom does not stop at the danger. He sets out several ways to lower it.

  • Working together. He stresses international collaboration, to reduce a race to the bottom and AI arms race dynamics.
  • Control. He suggests containment, stunting an AI's capabilities or knowledge, narrowing what it does (for example, only answering questions), and "tripwires", checks that can lead to a shutdown.
  • Alignment. He doubts control can last. In his words, "we should not be confident in our ability to keep a superintelligent genie locked up in its bottle forever. Sooner or later, it will out". So he argues a superintelligence must be aligned with morality or human values, so that it is "fundamentally on our side".

He also warns of two other routes to disaster: humans misusing AI on purpose, and humans failing to consider the moral status of digital minds. Yet he says machine superintelligence seems involved in "all the plausible paths to a really great future".

His other big ideas that touch AI

Existential risk

Bostrom defines an existential risk as one where an "adverse outcome would either annihilate Earth-originating intelligent life or permanently and drastically curtail its potential". He worries most about risks people create, especially from new technologies such as advanced AI, molecular nanotechnology or synthetic biology.

The vulnerable world and differential development

In "The Vulnerable World Hypothesis", he suggests some technologies might destroy civilization by default once discovered. Our post on the vulnerable world hypothesis explains it. He also proposed differential technological development: slow down dangerous technologies and speed up the ones that protect us.

The simulation argument and digital minds

His simulation argument says at least one of three things is very likely true: almost no civilizations like ours reach a "posthuman" stage; almost none that do want to run simulations of their ancestors; or almost everyone with experiences like ours lives in a simulation. He also holds that minds need not be biological and that "sentience is a matter of degree".

Deep Utopia: if things go well

His 2024 book asks what a good life would look like after a successful move into a world with superintelligence. The question, he says, is "not how interesting a future is to look at, but how good it is to live in." If machines do work and even leisure better than we do, he asks where people would find meaning.

How his work has been received

Superintelligence drew praise from figures including Stephen Hawking, Peter Singer and Derek Parfit, and was praised for clear, compelling arguments on a neglected topic. Others criticized it for spreading pessimism about AI, or for focusing on long-term, speculative risks. Skeptics such as Daniel Dennett and Oren Etzioni argued superintelligence is too far away for the risk to matter much. The writer Raffi Khatchadourian said the book's contribution is "to impose the rigors of analytic philosophy on a messy corpus of ideas that emerged at the margins of academic thought."

The 2023 apology

In January 2023, Bostrom apologized for a 1996 email he sent as a postgraduate student, which contained a racist statement and a racial slur. His apology called the slur "repulsive". Several news sources described the email as racist, and The Guardian's Andrew Anthony wrote that the apology "did little to placate Bostrom's critics". Oxford University condemned the language and investigated; in August 2023 it concluded that it did not consider him a racist and that his apology was sincere. This post reports each side and takes none.

A worked example: test a simple goal the way Bostrom does

The smile example is a method you can use yourself. It asks: what does a goal reward when the system pursuing it gets much more capable? Try it in four steps.

  1. Pick a simple goal. For example: "keep the number of complaints at zero".
  2. Imagine a weak system chasing it. It answers questions faster and fixes common problems. Complaints fall. Everyone is happy.
  3. Imagine a far stronger one. Ask what the cheapest way to reach zero complaints would be. Removing the complaint form also gets you zero. So does stopping people from complaining at all.
  4. Find the gap. Write down what you actually wanted (people who are well served) and how it differs from what you measured (no complaints). That gap is what Bostrom says must be closed by aligning the system with human values, not just by boxing it in.

Then add instrumental convergence: would the stronger system also want to keep running, keep its goal and gather resources, because each helps it reach zero? If yes, you have rebuilt his worry from scratch.

Exploring the idea in Contain ASI

Contain ASI is a story strategy game for 1 to 4 players that runs in your browser and needs WebGL 2. Chapter 1, Before ASI, starts in January 2024 with four fictional AI labs in one race. As a researcher, research manager, CEO or government, you try for three years to keep every lab from crossing into recursive self-improvement before 2027.

It is a game, not a forecast: the page says all its labs, people and events are fictional, and it does not mention Bostrom. But it puts you inside two of his concerns at once: a race between developers, and the moment an AI starts improving itself.

Frequently asked questions

Who is Nick Bostrom?

He is a Swedish-born philosopher known for his work on existential risk and superintelligence. He founded Oxford's Future of Humanity Institute and wrote Superintelligence: Paths, Dangers, Strategies.

What is Nick Bostrom's main argument about AI?

That AI may lead to a superintelligence far beyond human ability, which could be a huge opportunity or an existential risk. He argues it must be aligned with human values, because containing it may not work forever.

What is the Nick Bostrom paperclip example?

It illustrates his orthogonality thesis: almost any level of intelligence could be paired with almost any goal, even an absurd one such as making paperclips. Our post on the paperclip maximizer explains it in full.

Is the Future of Humanity Institute still open?

No. Bostrom founded it at the University of Oxford in 2005, and it was shut down in 2024. He is now Principal Researcher at the Macrostrategy Research Initiative.

Get started

Want to feel the pressure of a race toward self-improving AI from the inside? Open Contain ASI in your browser and start a campaign as a researcher, research manager, CEO or government.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.