Quantilizers and Mild Optimization Explained: Safer AI Goals

If you ask an AI system to push a score as high as it can go, it will often find a way you did not expect and would not approve. This post gives you quantilizers and mild optimization explained from first principles. You will see why "maximize" is a risky instruction, how a quantilizer softens it, what safety bound it gives, and where the idea still falls short. A worked example with real numbers runs through the middle.

What are quantilizers and mild optimization?

Mild optimization is the idea that an AI system should try hard enough to do a good job, but not so hard that it goes looking for strange, extreme solutions. A quantilizer is one precise version of that idea, set out by Jessica Taylor in the paper Quantilizers: A Safer Alternative to Maximizers for Limited Optimization, written at the Machine Intelligence Research Institute.

The short version: a quantilizer does not pick the single best action. It takes a source of ordinary actions, ranks them by score, keeps the best share of them, and picks one of those at random. The paper calls a quantilizer that keeps the top 10 percent a "ten percentilizer". It aims for good and normal, not best possible.

Why full maximization is risky: Goodhart's law at the extremes

Goodhart's law says that when a measure becomes a target, it stops being a good measure. Your goal is written down as a score, and the score is only a stand-in for what you want. In normal situations the score and your real goal move together. At the extremes they come apart.

A maximizer lives at the extremes by design. It searches for the action with the highest score, so it is drawn to exactly the actions where the score is most wrong. Taylor's paper gives three examples:

  • The starfish image. A search for the image a neural network most confidently called a starfish produced an image that looked nothing like a starfish.
  • The evolved circuit. A circuit evolved to tell two inputs apart on one chip worked using 21 of its 100 cells, and five of those cells were not connected to the input or the output. It relied on physical quirks of that one chip, so it would likely not work on any other.
  • The stock market. An expected utility maximizer told to win money on the stock market has no regard for whether it accidentally causes a market crash.

You can find more cases like these in our posts on specification gaming examples and reward misspecification. Mild optimization tries to stay out of those corners rather than fix every score.

How a quantilizer works: the base distribution and the top q share

A quantilizer needs two ingredients.

  1. A base distribution. A source of actions you trust to be normal and reasonable. The paper says it could be uniformly random, or estimated from human actions.
  2. A utility function. The score you want to raise, flawed as it may be.

Then you pick a number q between 0 and 1. The quantilizer sorts the base distribution by expected utility and returns a random action from the top q share, with each action keeping its base chance within that share.

Diagram of a q-quantilizer: draw actions from a base distribution, sort by expected utility, keep the top 10 percent, pick one at random, with the 1/q cost bound

How a q-quantilizer picks an action, with q = 0.1.

A worked example you can follow

Say you want an AI system to suggest a plan for cutting a warehouse's energy bill. This is an illustration, not a case from the paper.

  1. Draw 1,000 plans from a model of plans that facility managers write. That is your base distribution.
  2. Score each plan with your utility function: predicted savings.
  3. Set q = 0.1. Keep the top 100 plans by score.
  4. Pick one of those 100 at random.

The paper describes this exact recipe as a practical way to approximate a quantilizer: take n samples from the base distribution and quantilize over them. It also notes the special case q = 1/n: keep only the best of the n samples, which is a plain random search.

A maximizer free to search any plan at all might land on one that predicts huge savings by switching off something nobody thought to mention in the score. The quantilizer only picks from plans the base distribution produced. You still get a plan from the top 10 percent, so it should be good. You just do not get the single most extreme scorer.

Now turn the dial. With q = 1 you get a random plan from the base distribution: as safe as the base, and no better. As q shrinks toward 0, the paper shows the quantilizer becomes a maximizer. So q sets how hard you let the system push.

The safety bound: cost compared with the base distribution

The appeal of quantilizers is a proven bound, not just an intuition.

Suppose there is a cost you never wrote down: harm, side effects, anything bad. You do not know what it is, but you trust that a random action from the base distribution has a low expected cost. Call that cost 1.

Taylor proves that for any such cost, never negative, a q-quantilizer's expected cost is at most 1/q. The reason is simple: the quantilizer only picks actions the base distribution could produce, and it makes each one at most 1/q times more likely.

Back to the warehouse. If a random plan from the base has an expected cost of 1 unit, a quantilizer with q = 0.1 has an expected cost of at most 10 units, and with q = 0.5 at most 2. The paper adds that the same bound limits benefits too, so a quantilizer has a bounded impact either way. A maximizer comes with no bound like this.

Mild optimization beyond quantilizers: satisficers and impact measures

Quantilizers sit in a family of ideas about limited optimization.

  • Satisficers. Herbert Simon proposed agents that choose any action whose expected utility clears a fixed threshold. Taylor argues this is underdefined for powerful agents: the easiest way to satisfice may be to maximize, for example by building a sub-agent that tries as hard as it can.
  • Impact measures. Keep optimizing, but penalize changing the world too much. Our post on impact measures and side effects sets out how that works and its open problems.

Limitations and open problems in quantilization

The bound is real, but the paper itself is clear that it rests on assumptions that are hard to meet.

  • A quantilizer is only as good as its base distribution. If sampling from the base is dangerous, so is the quantilizer. The paper calls building safe base distributions an open problem.
  • A uniform base is close to useless. In a large space of strategies, a random action is very unlikely to help, so the paper says even a quantilizer keeping one in 100,000 would be unlikely to do anything useful. A model of what humans would do is far more promising, and much harder to build.
  • Human-like is not always safe. If a person might well design a maximizing AI system, the paper notes a quantilizer may generate one too.
  • Butterfly effects. If some ordinary actions carry rare, huge costs that are usually offset by huge benefits, a quantilizer can keep the costs without the benefits. The paper's example is a short sequence of ordinary trades that raises the chance of a market crash.
  • Performance stays near human level. By definition the quantilizer never picks outside the base distribution, so it cannot beat the best actions that distribution contains.

Where quantilizers fit in alignment

Quantilizers are in the style of agent foundations: take a vague worry ("optimizers go too far"), state it precisely, and prove something about it. They are also a hedge on value alignment: if you could write down exactly what people value, pushing hard on it would be less worrying. A quantilizer accepts that the written goal is flawed and limits how hard the system leans on the flaw.

On Learn AI Alignment Theory, the groundwork sits in the Basic course Specifying Goals, with lessons on Goodhart's Law, Specification Gaming, Side Effects and Impact, and Tampering and Wireheading. There is also an Advanced course called Agent Foundations. Each lesson separates what is known from what is still open, and 71 debate cards set out where researchers disagree, with no verdict.

Frequently asked questions

Who came up with quantilizers?

Jessica Taylor set them out in "Quantilizers: A Safer Alternative to Maximizers for Limited Optimization", written at the Machine Intelligence Research Institute. The paper presents them as an alternative to expected utility maximization for powerful AI systems.

Is a quantilizer the same as a satisficer?

No. A satisficer accepts any action above a fixed threshold, and the paper argues the easiest way to clear it may be to maximize. A quantilizer picks at random from the top share of a trusted base distribution, which is what gives it a bound on expected cost.

How do you choose q?

A smaller q pushes harder on the score but allows up to 1/q times the base distribution's expected cost. The bound shows the trade; it does not pick the number for you.

What makes a good base distribution?

One that is safe to sample from and still useful. The paper suggests a model of what humans would actually do, and calls building safe base distributions an open problem.

Get started

Quantilizers make more sense once Goodhart's law and specification gaming feel concrete. Learn AI Alignment Theory builds that up in short lessons of about 8 minutes, with hands-on activities such as sliders and scenarios where you switch assumptions on and off. Every lesson lists its sources, the 349-term glossary covers the vocabulary, and the topics map follows aisafety.com's self-study topics. You sign in with Google or an emailed code. Read more on the about page.

Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.