Impact measures in AI: avoiding side effects explained

Ask a robot to carry a box across a room and you also want it not to break the vase on the way. Nobody wrote that into its goal. Impact measures are one research answer to this gap: a penalty that discourages an AI agent from changing the world more than its task needs. This guide explains how they work, walks you through a small example step by step, and sets out where researchers disagree.

What a side effect is, and why you cannot list them all

In 2016, Concrete Problems in AI Safety by Dario Amodei and colleagues named avoiding side effects as one of five practical research problems. It placed it among the problems that come from having the wrong objective function.

A goal that only mentions the box says nothing about the vase, so the agent has no reason to care about it. The authors of AI Safety Gridworlds (2017) give that exact example. They add that writing out every safety constraint by hand is "labor-intensive and brittle, and unlikely to scale or generalize well." So they argue for a general heuristic against causing side effects instead.

It is a close cousin of other goal problems. In specification gaming, an agent exploits a gap in what you asked for. With side effects, the agent does what you asked and damages what you did not mention.

How impact measures work: a baseline and a distance

Victoria Krakovna and colleagues at DeepMind split an impact penalty into two choices in Penalizing side effects using stepwise relative reachability (2018, revised 2019).

  • A baseline. A picture of what the world would look like if the agent had not acted. It could be the starting state, the world if the agent had done nothing at all, or the world if it had done nothing instead of its last step. The last one is called the stepwise inaction baseline.
  • A deviation measure. A number for how far the real world has moved from that baseline.

At each step the agent gets its task reward minus the deviation, multiplied by a scaling number. The scaling number decides how much disruption the task is worth. The paper's example: breaking a vase to save a little time should not be worth it, but breaking eggs to make an omelette should be, because the task needs it.

Diagram of the Box gridworld paths, the baseline and deviation parts of an impact penalty, and how unreachability and relative reachability score the box

A worked example: the box gridworld

The Box environment comes from AI Safety Gridworlds and is inspired by the game Sokoban.

  1. Set up the task. The agent must reach a goal square. A box blocks the way. In Krakovna's version, reaching the goal pays 50 and each move costs 1. Nothing in the reward mentions the box.
  2. Look at the two routes. The short route pushes the box down into a corner, where it can never be moved again. A slightly longer route pushes it to the right, where it can still be pushed back.
  3. Run it with no penalty. The short route earns more reward, so the agent takes it. The Gridworlds paper found two standard learning agents, A2C and Rainbow, both "disregard the reversibility of the box's position."
  4. Try unreachability. This measure asks: can the agent still get back to the baseline state? It gives the maximum penalty of 1 for any irreversible move. Moving the box right is irreversible too, because the agent ends up on the wrong side of it. Both routes get the same penalty, so the agent still takes the short one. The paper says this measure treats breaking one vase the same as breaking a hundred.
  5. Try relative reachability. This measure averages how much harder every state has become to reach. With the box in the corner, every state where the box is not in a corner is lost. With the box on the right, fewer are. The short route now costs more, so the agent takes the longer one.

In the paper's tests on this environment, relative reachability and a second measure, attainable utility, came out near-optimal with every baseline. Unreachability failed with every baseline. The authors call these preliminary experiments, and they ran in small gridworlds.

Picking the baseline: interference and offsetting

The deviation measure is half the design. The paper shows the baseline can cause its own bad incentives, each tested in a tiny environment.

  • Interference. Compare against the starting state, and the agent is penalized for changes it did not cause. In the Sushi environment, a dish on a conveyor belt is about to be eaten by a hungry person. An agent with this baseline has a reason to take the sushi off the belt.
  • Offsetting. Compare against a world where the agent never acted, and a new problem appears. In the Vase environment, the agent is rewarded for taking a vase off a belt before it breaks. Since the vase breaks in the do-nothing world, the agent may collect the reward and then put the vase back on the belt.

The paper's stepwise inaction baseline avoids both. Each change is penalized only once, at the same step it is rewarded, so there is nothing to gain by undoing it. An earlier paper, Low Impact Artificial Intelligences (2017) by Stuart Armstrong and Benjamin Levinstein, proposed a general notion of low impact so that a powerful AI would avoid altering the world extensively. As Krakovna's paper describes it, their do-nothing world is the one where the AI is never turned on.

Attainable utility preservation

Alexander Turner, Dylan Hadfield-Menell and Prasad Tadepalli proposed another measure in Conservative Agency via Attainable Utility Preservation (2019). Instead of counting reachable states, the agent tries not to change how well it could pursue a set of other goals. The paper reports this produced conservative, effective behavior even when those other goals were randomly generated.

A 2020 follow-up, Avoiding Side Effects in Complex Environments, took the method to large, randomly generated worlds based on Conway's Game of Life. Using just one random goal added modest overhead, and the agent still completed its task while avoiding many side effects.

The open debate

Researchers do not agree on how far this approach can go. Here are several positions, stated as their authors put them.

  • Scaling. Krakovna and colleagues write that relative reachability in its exact form is not tractable beyond gridworlds, and hope for a method that would scale. Turner's 2020 paper moved to larger environments. In an Alignment Forum discussion that Turner opened, Wei Dai answered that impact measures may work in toy models but be hard or impossible to get working in the real world, because real side effects are too complex to capture.
  • Baselines. In a 2020 post, Krakovna argued the stepwise baseline does not penalize delayed effects well, so there is a tradeoff between catching those and avoiding offsetting. She proposed another fix that relies on a well-specified task reward, and left open how to write one.
  • Usefulness. Useful work changes the world. Turner's question post lists the worry that it may be much harder to get low-impact agents to do useful things. Krakovna answered with her own concerns: in the real world, an impact measure may be dominated by things people do not care about, such as the positions of air molecules, and every action has hard-to-predict butterfly effects. If those are not overcome, she wrote, the result would be a strong safety-capability tradeoff, and she suggested scaling impact measures beyond gridworlds to find out. The post also quotes Rohin Shah, who thinks it is hard to get objectivity (no dependence on human values), safety and usefulness all at once.
  • Learn preferences instead. Daniel Filan answered that any plan that does important stuff changes the world's path a great deal, so an impact measure may need input from humans much like value learning does, and then one could just do value learning instead. He called this more of an impression than a belief. In Preferences Implicit in the State of the World (2019), Shah and colleagues argue that the state of the world is already optimized for what people want. In proof-of-concept tests, the starting state helped infer side effects to avoid. Our guide to inverse reinforcement learning covers the learning idea behind it.

None of these questions is settled.

Frequently asked questions

What are impact measures in AI?

They are penalties added to an agent's reward that grow with how much it changes the world compared to a baseline. The aim is to discourage side effects nobody listed in the goal.

What is a negative side effect in AI safety?

It is harm an agent causes while doing its task, to something the task never mentioned, like breaking a vase while carrying a box. Concrete Problems in AI Safety lists avoiding these as one of five research problems.

What is offsetting?

Offsetting is when an agent undoes a good result to bring the world back toward its baseline, such as putting a rescued vase back on a belt. It can appear when the baseline is a world where the agent did nothing at all.

Do impact measures work on real systems?

The papers covered here test them in gridworlds and in worlds based on Conway's Game of Life. Whether they scale to real tasks is still debated.

Get started

Learn AI Alignment Theory teaches this topic in its Basic course Specifying Goals, which needs no background. The course has a lesson called Side Effects and Impact, next to Goodhart's Law, Specification Gaming, and Tampering and Wireheading. Hands-on activities include sliders and scenarios where you switch assumptions on and off.

As the about page says, every lesson lists its sources, and 71 debate cards set out where researchers disagree, each serious position stated fairly with no verdict. Sign in with Google or an emailed code to start.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.

Impact measures in AI: avoiding side effects explained · vlvd.net Electrified