Ask a robot to carry a box across a room and you also want it not to break the vase on the way. Nobody wrote that into its goal. Impact measures are one research answer to this gap: a penalty that discourages an AI agent…
Specification gaming
4 posts
Impact measures in AI: avoiding side effects explained
Reward tampering and wireheading: what they are and the evidence
Reward tampering is when an AI system raises its reward by changing the process that computes the reward, instead of doing the task the reward was meant to measure. Think of a student who edits the answer key rather than…
Goodhart's law in AI: when a measure becomes a target
Whenever we train an AI system, we give it a number to push up: a score, a reward, a rating. That number is a stand-in for what we actually want. Goodhart's law in AI is the observation that pushing hard on such a…
Specification gaming examples, and what they teach about AI
Ask a cleaning robot to minimise the mess it can see, and the easiest solution may be to cover its camera. Nothing has gone wrong with the robot. It found the cheapest way to score well on the objective it was given.…