Inverse reinforcement learning explained simply

Inverse reinforcement learning is a way to work out what someone wants by watching what they do. Ordinary reinforcement learning starts with a goal and learns behavior. Inverse reinforcement learning, or IRL, runs the other way: it starts with behavior and tries to recover the goal behind it.

This guide explains the idea in plain words, walks through a small example you can do on paper, and sets out why some researchers see IRL as a route to aligned AI while others think behavior alone can never be enough.

What inverse reinforcement learning means

In reinforcement learning, you hand an agent a reward: a number that says how good each outcome is. The agent tries things and keeps what scores well. The reward is the goal, written down by a person.

Saurabh Arora and Prashant Doshi, in A Survey of Inverse Reinforcement Learning (2018), define IRL as the problem of inferring the reward function of an agent, given its policy or observed behavior. A policy is simply what the agent does in each situation. So IRL asks one question: given what this agent did, which reward would make that a sensible thing to do?

Diagram comparing reinforcement learning, which goes from a reward to behavior, with inverse reinforcement learning, which goes from recorded behavior back to a guess at the reward

Why alignment researchers care

Writing a reward by hand is hard. Our guide to specification gaming examples shows what happens when the written goal misses what people meant: the system scores well and does the wrong thing.

IRL offers a different route. Instead of writing the goal, let the system learn it from people. If a person's choices reveal what they value, a system that reads those choices well might pursue the same values, including the parts nobody thought to write down.

It is one of several ways to learn from humans. RLHF learns from people comparing two answers. IRL learns from what people actually do.

A worked example: inferring a goal by hand

You can do the core move of IRL on paper. The scene below is made up for illustration.

  1. Watch the behavior. A delivery driver has two routes to the same street. Route A takes 10 minutes and passes a school. Route B takes 14 minutes and avoids it. Over a month, you see the driver take Route B on school days and Route A at weekends.
  2. List candidate rewards. Reward 1: arrive as fast as possible. Reward 2: arrive fast, but avoid busy school streets. Reward 3: pass as few children as possible, whatever the time.
  3. Test each one against what you saw. Reward 1 predicts Route A every day, so it does not fit the school days. Reward 3 predicts Route B every day, so it does not fit the weekends. Reward 2 predicts exactly the pattern you saw.
  4. Keep the best fit, and stay honest about it. Reward 2 explains the data. So might others: perhaps the driver dislikes the traffic near the school, not the risk to children. Both rewards predict the same routes.
  5. Ask what would tell them apart. A school holiday with no traffic would help. If the driver still avoids the school, the risk explanation gains ground.

Step 4 is the heart of the difficulty. Arora and Doshi list it among the central challenges of IRL: accurate inference is hard, and the answer is sensitive to what you assumed before you started.

The ambiguity problem

The example had a hidden assumption: that the driver chooses well, given what they want. Real people do not always. They get tired, misjudge, and follow habits.

Stuart Armstrong and Sören Mindermann take this on in Occam's razor is insufficient to infer the preferences of irrational agents (2017). They show that behavior cannot be uniquely split into a way of planning and a reward. For example, a person who reaches for cake could value cake, or could value health and be bad at acting on it. Preferring the simplest explanation does not reliably pick the true one either, they argue.

Their conclusion is that IRL needs some normative assumptions: claims about how a person's behavior ought to be read, which cannot be deduced from watching alone.

Cooperative inverse reinforcement learning

Classic IRL assumes the person acts well on their own while the system watches. Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell changed the setup in Cooperative Inverse Reinforcement Learning (2016).

In cooperative IRL, or CIRL, the human and the robot play a game together. Both are rewarded by the human's reward, but the robot does not know at first what that reward is. The authors use this game as a formal definition of the value alignment problem.

They show that good solutions to the game include active teaching by the human, active learning by the robot, and actions meant to communicate. These work better for alignment than a person simply performing well in isolation. Once the robot's uncertainty is part of the problem, teaching and learning become part of the answer.

Where researchers disagree

  • The hopeful view. The CIRL authors present learning a person's reward inside a cooperative game as a formal way to define and work on value alignment. On this view, the system's uncertainty about the goal is what turns teaching, asking and learning into part of the solution.
  • The skeptical view. Armstrong and Mindermann argue that no amount of watching settles what someone values unless you add assumptions about how rational they are. On this view, the hard part of alignment moves into those assumptions rather than going away.
  • The practical view. The Arora and Doshi survey lists limits that matter today: the cost of solving IRL grows fast with the size of the problem, and the result depends on prior knowledge and on how well the learned reward carries over to new situations.

These positions are not settled. Learn AI Alignment Theory sets them out side by side, with no verdict.

How the courses cover inverse reinforcement learning

Learn AI Alignment Theory teaches this in its Intermediate course Learning from Humans. Its lessons are Inferring Goals from Behavior, Learning Rewards from Comparisons, Constitutional AI and the Limits of Feedback, and The Theory of Reward Learning. The Intermediate level is about how today's training methods can go right or wrong.

The Three levels panel on the Learn AI Alignment Theory about page, showing Basic with 5 courses, Intermediate with 10 courses and Advanced with 7 courses

Every lesson lists its sources, so you can read papers like the ones above next. For another way of learning from people, our guide to Constitutional AI covers training with written principles instead. The about page shows the three levels and what you do in each lesson.

Frequently asked questions

What is inverse reinforcement learning in simple terms?

It is working backwards from behavior to the goal behind it. You watch what an agent does and infer which reward would make those choices sensible.

How is IRL different from copying behavior?

Copying repeats the actions. IRL tries to recover the reason for them, which is what a system would need to act sensibly in situations the person never showed.

Why can't IRL just find the true reward?

Many rewards can explain the same behavior, and people do not always act on what they want. Armstrong and Mindermann argue that extra assumptions about rationality are needed to choose between explanations.

What does the cooperative version add?

In cooperative IRL the person and the system share the person's goal, and the system starts out unsure of it. That makes teaching and learning part of the solution.

Get started

Start Learn AI Alignment Theory: sign in with Google or an emailed code, work through the Basic level, then take Learning from Humans for inferring goals, reward learning and the limits of feedback, with every side of the debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.