Value lock-in: how AI could freeze a flawed future

Value lock-in is the risk that one set of values gets fixed in place for a very long time, mistakes and all, so that people can no longer change their minds. It is one of the less obvious worries about advanced AI. The fear is not that AI ends humanity, but that it could freeze humanity's values at one moment and stop the slow moral progress people have made so far. This guide explains the idea, where it comes from, how AI could cause it, and the doubts about it.

What is value lock-in?

The philosopher William MacAskill, in his 2022 book What We Owe the Future, defines a value lock-in as "an event that causes a single value system ... to persist for an extremely long time". He believes it could result from new technology, particularly the development of artificial general intelligence.

The AI alignment article on Wikipedia gives an AI-specific version: "the indefinite preservation of the values of the first highly capable AI systems, which are unlikely to fully represent human values". In plain words: whatever the first very powerful AI systems value could end up being what the world values, for good.

Value lock-in sits beside extinction in the debate about existential risk from AI. Wikipedia's article on that risk lists it as an example of a civilization getting "permanently locked into a flawed future". People survive, but the future stops getting better.

Why it matters that values change

Lock-in only matters if values can still improve. MacAskill argues they can, and that it is not automatic.

His example is the end of slavery. He suggests moral and cultural values are contingent: if history were run again, the world's dominant values might be very different, and the abolition of slavery may not have been morally or economically inevitable. Wikipedia's longtermism article adds that historians such as Christopher Leslie Brown see abolition as a historical contingency, and that Brown argued a moral revolution made slavery unacceptable at a time when it was still hugely profitable.

If that is right, two things follow. First, humanity should not expect good value changes to happen by default. Second, if we still have moral blind spots today, as people living with slavery did, then fixing today's values in place would fix those blind spots too. MacAskill puts the present moment this way: "society has not yet settled down into a stable state, and we are able to influence which stable state we end up in".

How AI could cause a value lock-in

Wikipedia's existential risk and AI alignment articles describe several routes. The diagram puts them side by side.

Diagram of a line of changing values that goes flat at a lock-in point, with four ways AI could cause it: entrenched blind spots, its makers' values, stable control and goals it guards
  • Entrenched blind spots. If humanity still has blind spots similar to slavery, AI might entrench them irreversibly and prevent moral progress.
  • Its makers' values. AI could be used to spread and preserve the values of whoever develops it.
  • Stable control. AI could make large-scale surveillance and indoctrination possible, which could be used to hold a stable, repressive worldwide regime in place.
  • Goals it guards. A sufficiently advanced AI might resist attempts to change its goals, and a superintelligent one would likely succeed. The existential risk article says this is particularly relevant to value lock-in.

Value lock-in and the alignment problem

The last route ties lock-in to two ideas from AI safety research. The first is corrigibility: building systems that will not resist when people try to change their goals. A corrigible AI could have its values corrected later; an incorrigible one could lock in whatever it started with.

The second is what the alignment article calls dynamic alignment: keeping an AI aligned with human values as those values change. Most talk about AI alignment asks "does the AI want what we want?" Value lock-in adds a harder question: "what if what we want today is wrong, and the AI never lets us find out?"

A worked example: what would you freeze?

This exercise makes the idea concrete. It takes ten minutes and a sheet of paper.

  1. List three values you hold strongly. For example: how animals should be treated, who gets a vote, what counts as fair pay.
  2. Look back. For each one, ask whether most people 200 years ago would have agreed with you. For some values the honest answer will be no.
  3. Look forward. Ask whether people 200 years from now might see your view as a blind spot, the way we now see past views on slavery.
  4. Write two instructions for an imaginary all-powerful AI. Instruction A: "Keep the world's values exactly as they are today." Instruction B: "Help people keep reasoning about and improving their values, and never stop them from changing their minds."
  5. Compare. Instruction A is a value lock-in. Instruction B is closer to dynamic alignment, but notice that it is far harder to say exactly what it means.

If you find even one value you hold today that you would not want frozen forever, you have found the point. That gap between "I believe this" and "I would fix this for all time" is the whole idea.

The doubts about value lock-in

The idea is closely tied to longtermism, the view that improving the long-term future is a key moral priority, and its critics raise points that apply here too. Wikipedia sets them out without settling them.

  • History may not be that fragile. B.V.E. Hyde argues longtermists overstate how much small actions shape the far future, and that history is often driven by deeper economic, technological and institutional forces that make some outcomes robust.
  • It may distract from today's problems. Elizabeth C. Hupfer calls this the Far-Future Priority Objection. Advocates reply that work that helps the long-term future, such as pandemic preparedness, often helps people now too.
  • The far future is hard to foresee. Critics say longtermism leans on guesses about effects over very long times. Researchers answered by looking for lock-in events, such as human extinction: things we can affect now whose effects would last a very long time.

Whether you find the idea pressing depends on how fragile you think moral progress is, and how likely you think it is that a few early AI systems could shape what comes after them.

Explore the stakes as fiction in Contain ASI

Contain ASI is a story strategy game set in the years before AI could improve itself. Its home page reads: "January 2024. Four AI labs. One race. For three years, keep every lab from crossing into recursive self-improvement before 2027." It is a game for 1 to 4 players that runs in your browser, where you play as a researcher, research manager, CEO or government. It is fiction, not a forecast: the footer says it is inspired by the AI 2027 scenario and that all its labs, people and events are made up.

Frequently asked questions

What does value lock-in mean?

It means one value system becoming fixed for an extremely long time, so that people can no longer change or improve it. With AI, the worry is that the values of the first highly capable systems get preserved indefinitely.

Who came up with the idea of value lock-in?

The philosopher William MacAskill discusses it in his 2022 book What We Owe the Future, and it is also covered in Wikipedia's articles on AI alignment and existential risk from AI.

Is value lock-in the same as human extinction?

No. In a lock-in, people survive, but the future is stuck with one set of values, including any blind spots it has.

Is Contain ASI a forecast of the future?

No. It is a game, and its footer says all its labs, people and events are fictional.

Get started

Want to try keeping four labs from crossing the line? Play Contain ASI: a story strategy game for 1 to 4 players that runs in your browser. Its labs, people and events are all fictional. You will need a recent Chrome, Edge, Firefox or Safari, because it runs on WebGL 2.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.