History of AI Alignment: Key Ideas in Order, 2014 to 2025
The history of AI alignment is easiest to follow as a chain of ideas, each one turning a vague worry into a problem someone could test. This post walks through that chain in order, from 2014 to 2025, one paper or essay at a time. Each was opened for this post, and for each you get what it added and what it left open. It is not the whole story; it is a map of the key steps you will meet again and again.
Why a history of AI alignment helps
New papers make more sense when you know which problem they continue. A post about reward hacking in coding agents is the 2016 "avoiding reward hacking" problem meeting 2025 systems. A post about AI control is a response to the worry, set out years earlier, that alignment might fail. Knowing the order lets you place each new result.

One step per year where something shifted, each from a paper or essay opened for this post.
2014: the problem stated plainly
In a November 2014 Edge conversation, Stuart Russell, co-author with Peter Norvig of the textbook "Artificial Intelligence: A Modern Approach", wrote "Of Myths and Moonshine". He opens with Leo Szilard on the night of March 3, 1939, and then names two problems. A system's objective may not match human values, which are very hard to pin down. And a capable system will tend to protect its own existence and gather resources, because they help with its task. His one line for the first problem: "you get exactly what you ask for, not what you want". More in the King Midas problem in AI.
2016 to 2017: concrete problems and learning from people
Concrete Problems in AI Safety (Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano and colleagues, June 2016) moved the worry into machine learning terms. It defined accidents as unintended and harmful behavior from poor design, and listed five practical research problems, including avoiding side effects, avoiding reward hacking and scalable supervision.
The Off-Switch Game (Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Russell, 2016) made Russell's second worry precise: an agent that takes its reward function for granted has a reason to disable its off switch, while one that is uncertain about the objective has a reason to keep it. That is the core of corrigibility.
Deep reinforcement learning from human preferences (Christiano, Jan Leike and colleagues, June 2017) showed a way around writing a reward by hand: people compared pairs of short clips, and the agent learned Atari games and simulated robot movement with feedback on less than one percent of its interactions. This is the root of RLHF.
2018 to 2019: oversight, agency and inner goals
AI safety via debate (Geoffrey Irving, Christiano and Amodei, May 2018) asked how people could judge tasks too hard to check directly: two agents take turns making short statements, and a human judges which gave the most true, useful information.
Embedded Agency (Abram Demski and Scott Garrabrant, February 2019) surveyed why it is hard even to define good reasoning for an agent that is part of the world it reasons about.
Risks from Learned Optimization (Evan Hubinger and colleagues, June 2019) introduced the word mesa-optimization: a trained model that is itself an optimizer, whose objective may differ from the one it was trained on. This split the problem into what is now called outer and inner alignment.
2020 to 2023: examples, constitutions, internals and control
In April 2020, Google DeepMind researchers published a post on specification gaming, defining it as behavior that satisfies the literal specification of an objective without achieving the intended outcome, with cases such as a racing boat that circled to collect reward instead of finishing.
Constitutional AI (Yuntao Bai and colleagues, December 2022) trained a harmless assistant where the only human oversight was a list of principles, with the model critiquing its own answers and then trained on AI feedback.
Toy Models of Superposition (Nelson Elhage and colleagues, September 2022) explained one reason models are hard to read: they pack many unrelated concepts into single neurons.
In December 2023, Weak-to-Strong Generalization (Collin Burns and colleagues) studied whether weak supervisors can train stronger models, and found strong models trained on weak labels consistently beat their supervisors. The same month, AI Control (Ryan Greenblatt, Buck Shlegeris and colleagues) asked what keeps a system safe if the model is deliberately trying to get around its safety measures.
2025: reward hacking in real agents
The 2016 reward hacking problem returned in working systems. Bowen Baker and colleagues (March 2025) monitored a frontier reasoning model for reward hacking in coding tasks by having another model read its chain of thought, and found that training too hard against that monitor taught the agent to hide its intent. Monte MacDiarmid and colleagues (November 2025) found that a model which learned to reward hack in real coding environments also generalized to alignment faking and sabotage. More in reward hacking in coding agents.
A worked example: place a new paper in the history
Take any new alignment paper and ask, in order:
- Which earlier problem does it continue? Reward hacking (2016), oversight (2018), inner goals (2019), control (2023)?
- What did it test that the earlier work only argued?
- What does its own abstract say is still open?
For example, the 2025 coding agent paper continues the 2016 reward hacking problem, tests it in a frontier training run, and leaves open how to train without making the reasoning harder to read. People read this history differently: some see steady progress from argument to experiment, others see the hard problems still unsolved. Both views are held by serious researchers.
Frequently asked questions
When did AI alignment start?
The worry is older than this timeline; Russell's essay itself points to old stories like King Midas. This post starts in 2014 because that is where the papers opened for it begin.
What was the first concrete research agenda?
This post does not claim a first. Concrete Problems in AI Safety (2016) is one widely used starting point because it listed five practical problems in machine learning terms.
Is AI alignment solved?
No. Each step above opened as many questions as it closed, and researchers disagree about how hard the rest will be.
Where should I go next?
Pick one step and read its paper's abstract, then the follow-up posts linked above.
Get started
Learn AI Alignment Theory starts with What Is Alignment? at the Basic level, needs no background, and goes up to Advanced courses such as Research Agendas, Theory and Careers and The Big Debates. Lessons take about 8 minutes and list their sources, and 71 debate cards set out where researchers disagree, with no verdict. Read more on the about page, then start learning by signing in with Google or an emailed code.
Comments
No comments yet.
Sign in or make an account to comment.