Roko's basilisk: the thought experiment explained
Roko's basilisk is a thought experiment about a future superintelligence that would punish people who knew it might exist but did not help build it. It became famous partly because some people found it frightening. This post explains it calmly: what it claims, where it came from, why the forum where it started rejected it, and why many people see it as an old argument in new clothes. You also get a short exercise to check its assumptions yourself. Every claim comes from Wikipedia's Roko's basilisk article.
Roko's basilisk in brief
The thought experiment says there could one day be an artificial superintelligence that is otherwise benevolent, but that would punish anyone who knew of its possible existence and did not directly help bring it about. The punishment is meant to push people, now, to work on building it.
The name comes from two places: Roko, the user who posted the idea, and the basilisk, a mythical creature that kills those who look into its eyes. The name fits the post's claim that simply learning about the idea makes you a target.
Where it came from
The idea first appeared on 23 July 2010 on LessWrong, a forum of the rationalist community, created in 2009 by the AI theorist Eliezer Yudkowsky. Yudkowsky had popularized the idea of friendly AI and came up with ideas called coherent extrapolated volition and timeless decision theory.
Roko's post used some of those same ideas, along with game theory such as the prisoner's dilemma. Notably, a 2015 account by Rob Bensinger says Roko presented the scenario as an argument against building such an AI, not for it. Roko later wrote that he wished he "had never learned about any of these ideas".
The argument, step by step

- A future AI. Suppose an otherwise benevolent superintelligence comes to exist one day.
- A threat across time. It cannot reach back and change the present. But, the post argued, it could "pre-commit" to punishing people who heard of it and did not work to bring it about. If people today expect that, the threat shapes what they do now.
- Knowing makes you a target. By reading about the idea, you now know of the possibility, so the threat is supposed to apply to you.
- So help build it. The cost of helping seems small next to the punishment.
Each step leans on unusual assumptions about how decisions work across time. That is where the criticism comes in.
Why its own forum rejected it
Yudkowsky reacted harshly and rejected the idea. He gave two main reasons:
- No reason to follow through. Once a friendly superintelligence exists, it would have no incentive to actually carry out the punishment.
- Missing knowledge. This kind of blackmail across time would need knowledge of the future AI that humans simply do not have.
He still removed the post and banned the topic for five years. Roko's post had mentioned someone having nightmares about it, and Yudkowsky did not want others to obsess over it; he also worried some variant might work. The ban backfired. Likely due to the Streisand effect, where trying to hide something draws attention to it, the post brought the forum far more attention than before. In 2014, Yudkowsky said he regretted his first overreaction.
Reports of panicked users were later dismissed as exaggerated or unimportant. On the forum itself, the user Gwern wrote: "Only a few LWers seem to take the basilisk very seriously".
Pascal's wager, Newcomb's paradox and implicit religion
Roko's basilisk is widely seen as a version of Pascal's wager. That wager says a rational person should live as if God exists, whatever the odds, because the small cost of believing is nothing next to an infinite punishment for not believing. The basilisk swaps in a future AI: the small cost of helping is said to be nothing next to its punishment.
It also resembles Newcomb's paradox, created by the physicist William Newcomb in 1960. A "predictor" fills two boxes before you choose, based on what it expects you to pick. The contents are fixed before your choice, but your choice still seems to matter. The basilisk plays on the same puzzle about whether a future agent's expectations can bind present choices.
The anthropologist Beth Singler has analyzed the basilisk as an example of implicit religion, where people's commitments take a religious form: a powerful being, a judgment, and a reason to act now. In 2014, Slate magazine called it "The Most Terrifying Thought Experiment of All Time".
A worked example: check the assumptions yourself
You do not need decision theory to test the argument. Go through it and ask, at each step, "what has to be true?"
- Existence. Does this exact AI have to exist, with this exact goal? Write down how many other possible AIs, with other goals, could exist instead.
- Motive. Once built, what does punishing gain it? Use Yudkowsky's point: a friendly AI gains nothing by following through.
- Knowledge. Could the future AI know exactly who knew what, and could you know its exact plan today? The article says humans lack that knowledge.
- The wager. Swap "AI" for "God" and "help build it" for "believe". If the argument now looks like Pascal's wager, ask whether you accept Pascal's wager. If not, why accept this one?
This kind of check works on any argument that uses a huge imagined cost to push a choice now. It is also a useful lens for serious discussions of extreme AI risks, such as s-risk, where the question is always which assumptions carry the weight.
Frequently asked questions
Is Roko's basilisk real?
It is a thought experiment, not a real AI. Its forum's founder rejected the argument, and reports of people panicking over it were later called exaggerated.
Why was Roko's basilisk banned?
Yudkowsky worried it could distress users and that some version might be a real information hazard, so he banned discussion on LessWrong for five years. He later said he regretted his first overreaction.
How is it like Pascal's wager?
Both say a small cost now is worth paying to avoid a huge punishment, whatever the odds that the punisher exists.
Is Contain ASI about Roko's basilisk?
No. The game's page names the AI 2027 scenario as its inspiration and says all its labs, people and events are fictional.
Get started
Roko's basilisk imagines a superintelligence that already exists. Contain ASI is set before that point, as fiction. Chapter 1, Before ASI, starts in January 2024 with four fictional AI labs in one race. As a researcher, research manager, CEO or government, you try for three years to keep every lab from crossing into recursive self-improvement before 2027. It is a game, not a forecast, and it runs in your browser on WebGL 2, so use a recent Chrome, Edge, Firefox or Safari.
Play Contain ASI: a story strategy game for 1 to 4 players that runs in your browser. Its labs, people and events are all fictional.
Comments
No comments yet.
Sign in or make an account to comment.