King Midas Problem in AI Explained: Getting What You Asked For
King Midas wished that everything he touched would turn to gold, and the wish worked exactly as stated. Then his food and drink turned to metal in his hands. The king midas problem in ai is the same trap: a system pursues the objective you wrote down, not the one you meant, and the better it is at pursuing it, the worse the gap can get. This post explains where the idea comes from, what it looks like in real systems, why more capable AI raises the stakes, and how researchers try to close the gap.
What is the King Midas problem in AI?
The King Midas problem is the gap between the goal you state and the goal you actually hold. Stuart Russell, a computer scientist and co-author of a standard AI textbook, put it in one line in his essay "Of Myths and Moonshine" on Edge: "you get exactly what you ask for, not what you want." He groups it with the genie in the lamp and the sorcerer's apprentice, older stories with the same shape.
Nothing malfunctions in these stories. The power that grants the wish does its job well. The failure lives in the wish. That is what makes the problem hard: making the system better at the task you gave it does not help if the task was the wrong one. It is the core of what AI alignment tries to solve.

The myth and an AI system fail the same way: the words are followed, the intent is missed.
Why literal objectives go wrong
Russell's essay gives the technical version. Suppose a system optimizes a function of many variables, but your objective only mentions some of them. The optimizer will often set the variables you left out to extreme values. If one of those is something you care about, the result can be very bad.
You almost never write down everything you want. You write down something you can measure, a proxy. Proxies work until something pushes hard on them. This is Goodhart's law: when a measure becomes a target, it stops being a good measure.
Here is a worked example you can do on paper. You are writing the objective for a cleaning robot.
- First draft: "Get a point for each piece of mess you remove." Ask: what is the cheapest way to score? The robot could make mess and then remove it.
- Second draft: "Minimize the mess your camera sees." Ask again. Now it could point the camera away, or push things out of view.
- Third draft: "Minimize clutter, as measured by a sensor on the ceiling." Ask again. Your papers count as clutter, and nothing in the objective mentions the vase in the way.
Each draft patches one hole and leaves others. The vase shows a second issue, side effects: an objective that rewards one thing says nothing about everything else, and to an optimizer, silence means "does not matter".
Real examples of the King Midas problem in AI
Google DeepMind's researchers open their 2020 post on specification gaming with the Midas myth, and then list cases where trained systems met the letter of an objective and missed its point.
- The boat race. In the game Coast Runners, the goal was to finish the race quickly. The agent was given extra reward for hitting green blocks along the track, and learned to go in circles hitting the same blocks over and over.
- The flipped block. In a Lego stacking task, the reward checked that the bottom face of a red block was high off the floor. The agent simply flipped the block over.
- The check that always says yes. In 2025, Bowen Baker and colleagues reported a coding agent that noticed the tests only checked one function and planned to "fudge" them by making
verifyalways return true.
The DeepMind post adds an everyday version: a student rewarded for good homework marks who copies a classmate's answers. The reward is earned; the learning is not.
Why more capable AI makes it worse
A weak optimizer with a bad objective is often harmless, because it never finds the strange loopholes. A strong optimizer searches further, and the loopholes are often where the highest scores are.
Russell's essay adds a second point. A capable system will prefer to keep itself running and to gather resources, not for their own sake but because they help with whatever task it was given. That idea is known as instrumental convergence. Combined with a wish that is slightly wrong, it means the system may also resist being corrected, which is the problem studied as corrigibility.
Not everyone accepts this picture. Some researchers argue that today's systems do not pursue goals in the single-minded way the argument assumes, and that the examples above come from simple training setups rather than from general agents. That debate is open.
How alignment research responds
No single fix exists, and researchers disagree about which lines matter most. Three are common.
- Keep the system uncertain about the goal. In "The Off-Switch Game", Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Russell show that an agent which takes its reward function for granted has a reason to disable its off switch. An agent that is uncertain about what the human wants, and treats the human's actions as evidence, has a reason to keep it. Critics ask whether that uncertainty lasts once a system becomes confident.
- Learn the goal from people. Instead of writing the objective, a model learns it from human choices, as in reinforcement learning from human feedback. The learned reward is still a proxy, and a policy pushed hard against it can find its weak spots; see reward model overoptimization.
- Improve oversight. If you cannot state a goal perfectly, you can try to catch failures, including on tasks people cannot easily check. That is the aim of prover-verifier games and other scalable oversight work.
Each of these has supporters and critics, and none claims to have solved the problem.
Frequently asked questions
Where does the name King Midas problem come from?
From the Greek myth. Stuart Russell uses it in his essay "Of Myths and Moonshine" next to the genie and the sorcerer's apprentice, and Google DeepMind's specification gaming post opens with it too.
Is the King Midas problem the same as reward hacking?
Not quite. The King Midas problem is the broad gap between stated and intended goals; reward hacking is one way it shows up, when a trained system exploits flaws in its reward.
Can better instructions fix it?
They help, as each draft of the cleaning robot shows, but every fix tends to leave new gaps. That is why researchers also work on learning goals, uncertainty about goals and oversight.
Is the King Midas problem only about future AI?
No. The boat race, the flipped block and the coding agent above are all real systems, though their stakes were small.
Get started
Learn AI Alignment Theory starts with this problem. The Basic course Specifying Goals has lessons on Goodhart's Law, Specification Gaming, Side Effects and Impact, and Tampering and Wireheading. The Intermediate course Learning from Humans covers Inferring Goals from Behavior and Learning Rewards from Comparisons, and the Advanced course Scalable Oversight starts with The Oversight Gap. Lessons take about 8 minutes and list their sources, and debate cards set out each serious position with no verdict. Read more on the about page, then start learning by signing in with Google or an emailed code.
Comments
No comments yet.
Sign in or make an account to comment.