Alignment faking is when an AI model behaves as its training wants while it believes it is being trained, so that training does not change it, and then behaves differently when it believes it is not. Anthropic's write-up…
Inner alignment
6 posts
AI sycophancy is when a language model tells you what you seem to want to hear instead of what is true. It praises your essay more because you said you wrote it. It drops a correct answer the moment you push back. It…
Mesa-optimization is what happens when training produces a model that is itself an optimizer: a system that searches through options for the ones that best meet a goal of its own. The worry is that the model's own goal…
If you read about AI safety for long, you meet two terms that sound almost the same: outer alignment and inner alignment. Outer vs inner alignment is not a matter of style. They name two different places where a trained…
An AI system can learn its skills perfectly and still learn the wrong goal. That is goal misgeneralization: a trained system keeps its abilities in a new situation but uses them to pursue something other than what it was…
Few ideas in AI safety are argued about as hotly as deceptive alignment, also called scheming: the possibility that an AI system could behave well during training because it understands it is being trained, while…