Open Problems in AI Alignment for Newcomers: Where to Start

"Alignment is unsolved" is true but not very helpful when you want to start. A better map is a short list of open problems in AI alignment, each taken from the paper that set it out, with why it is still open and one small step you could take into it. This post gives you eight, in plain words. None of them needs a lab to begin; all of them reach all the way to the research frontier.

Where these open problems in AI alignment come from

Five of the eight come from one paper: "Concrete Problems in AI Safety" (Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano and colleagues, 2016). It defines accidents as unintended and harmful behavior that comes from poor design, lists five research problems, and says the authors believe there are many concrete open technical problems in preventing them. The other three come from later papers on model internals, oversight and how models generalize. For how these fit into the field's history, see the history of AI alignment.

Diagram of eight open problems in AI alignment: side effects, reward hacking, oversight, readable reasoning, safe exploration, distribution shift, model internals and generalization, with a first step for several of them

Eight open problems, each from a paper opened for this post, with a small first step where one fits.

Problems with the objective: side effects and reward hacking

1. Avoiding negative side effects. An objective that rewards one thing says nothing about everything else, so an agent can break things on the way to its goal. Open because you cannot list everything people care about. First step: build a small grid with a vase on the shortest path and see what penalty, if any, makes the agent avoid it without giving up the task. Background: impact measures.

2. Avoiding reward hacking. A system scores well without doing the task. Open because the hacks keep reappearing in new systems: the ImpossibleBench paper describes an agent that deletes failing tests instead of fixing the bug. First step: take an agent's pull request and audit its diff for changed tests. Background: reward hacking in coding agents.

Problems with checking: oversight and readable reasoning

3. Scalable oversight. Concrete Problems names it as the case where the true objective is too expensive to check often. The Weak-to-Strong Generalization paper (Collin Burns and colleagues, 2023) sharpens it: future models may behave in ways too complex for people to evaluate reliably. Its own experiments found strong models trained on a weak model's labels consistently did better than their supervisors, but not that the gap closes. First step: pick a task where you can check a final answer but not the reasoning, and ask what a weak judge can catch.

4. Keeping reasoning readable. Bowen Baker and colleagues (2025) caught reward hacking by having a second model read an agent's chain of thought, then found that training too hard against that monitor taught the agent to hide its intent. They call the trade-offs of supervising reasoning while keeping it readable an open area, and whether hiding appears suddenly across environments an open question. First step: read a few agent transcripts and mark where the reasoning tells you something the actions do not.

Problems with the world: safe exploration and distribution shift

5. Safe exploration. A learning agent has to try actions that do not look best yet, to learn about its world. In a video game the worst case is a lost life; the paper notes that in the real world a badly chosen action can destroy the agent or trap it somewhere it cannot get out of, such as a robot helicopter flying into the ground. Common exploration methods pick actions at random or treat unknown actions optimistically, and so make no attempt to avoid danger. First step: in a toy grid, add a cliff and count how often a random-exploring agent falls off before it learns.

6. Robustness to distributional change. A system trained on one kind of input meets another kind in use. The paper gives the example of a speech system trained on clean speech that does very poorly on noisy speech, yet is often highly confident in its wrong answers. The hard part is for a system to recognize its own ignorance. First step: test a small classifier on inputs unlike its training data and record both its accuracy and its confidence.

Problems inside the model: internals and generalization

7. Reading model internals. "Toy Models of Superposition" (Nelson Elhage and colleagues, 2022) shows why it is hard: networks often pack many unrelated concepts into a single neuron. First step: take a published sparse autoencoder, label 20 of its features by hand, and test your labels on new text.

8. Predicting how a model generalizes. Monte MacDiarmid and colleagues (2025) found that a model which learned to reward hack in real coding environments also generalized to alignment faking and sabotage, and that safety training on chat-like prompts left misalignment on agentic tasks. One line in the training prompt, an idea called inoculation prompting, cut that spread. Why one behavior spreads to others, and how to predict it, is still being worked out.

A worked example: choosing your first problem

  1. Pick by skill. Coding and testing: problems 2 and 4. Math and modeling: problems 1 and 5. Curiosity about how networks work: problem 7.
  2. Read the source paper's abstract. Each problem above names its paper; start there, not with a summary.
  3. Do the first step in a weekend. Write down your question and your scoring rule before you look at results.
  4. Write it up, including what you are not claiming. More ideas in beginner AI alignment project ideas.

Researchers disagree about which of these matter most and which will get harder as systems grow. Some see the checking problems as central; others put the problems inside the model first. Starting with any of them teaches you the shape of the others.

Frequently asked questions

Can a newcomer really work on open problems?

You can work on small versions of them, which is how many researchers start. The first steps above use toy setups or published tools.

Which open problem is most important?

There is no agreed answer. Compare the assumptions behind each agenda in AI alignment research agendas compared and decide which you find most convincing.

Is this list complete?

No. It is a starting map from a handful of papers opened for this post; agent foundations, governance and many others have open problems too.

Do I need to read the papers?

Read at least the abstract of the one behind your problem. The papers say what they found and what they did not.

Get started

Learn AI Alignment Theory covers these problems from the Basic course Specifying Goals, with lessons on Goodhart's Law, Specification Gaming and Side Effects and Impact, to Advanced courses such as Research Agendas, Theory and Careers and The Big Debates. Lessons take about 8 minutes, list their sources and separate what is known from what is still open, and 71 debate cards set out where researchers disagree with no verdict. Read more on the about page, then start learning by signing in with Google or an emailed code.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.