Untrusted monitoring is a safety technique from the field of AI control: you use one copy of a powerful AI model to check the work of another copy, and assume either copy might be trying to trick you. It sounds…
AI alignment
35 posts
Linear probes are one of the simplest tools for looking inside an AI model. A probe is a small classifier that tries to read one property, such as "is this statement true?", straight from the numbers inside a network. If…
Eliciting latent knowledge, usually shortened to ELK, is the problem of getting an AI to tell you what it actually knows, rather than what it predicts you would believe. That sounds like a small difference. The…
Alignment research careers are open to more people than most guides suggest, but the roles differ a lot in what they ask of you. Some need a PhD and strong research taste, some need strong programming, and many need…
Direct preference optimization, or DPO, is a way to teach a language model which answers people prefer without building a separate reward model and without reinforcement learning. You give it pairs of answers to the same…
Ask a robot to carry a box across a room and you also want it not to break the vase on the way. Nobody wrote that into its goal. Impact measures are one research answer to this gap: a penalty that discourages an AI agent…
AI sandbagging is when an AI model does worse on a test than it really can, on purpose. Labs use test scores to decide whether a model is safe to release, and those scores are becoming part of AI regulation. If a model…
Alignment faking is when an AI model behaves as its training wants while it believes it is being trained, so that training does not change it, and then behaves differently when it believes it is not. Anthropic's write-up…
Reward model overoptimization is what happens when you train an AI model too hard against a learned reward model: the reward model's score keeps rising, but the quality you actually wanted levels off and then falls. It…
AI sycophancy is when a language model tells you what you seem to want to hear instead of what is true. It praises your essay more because you said you wrote it. It drops a correct answer the moment you push back. It…
Embedded agency is the problem of how an agent should reason and act when it is part of the world it is trying to change. Most theories of rational choice quietly assume the agent sits outside its world, like a player…
Reward tampering is when an AI system raises its reward by changing the process that computes the reward, instead of doing the task the reward was meant to measure. Think of a student who edits the answer key rather than…
AI jailbreaks are inputs that get a safety-trained language model to do what its safety training was meant to stop. A model is trained to refuse certain requests, and a jailbreak finds a way around that refusal. Studying…
If you are comparing AI alignment vs AI safety, here is the short answer most people mean: alignment is about whether an AI system pursues what its designers intend, and safety is the wider effort to stop AI from causing…
Mesa-optimization is what happens when training produces a model that is itself an optimizer: a system that searches through options for the ones that best meet a goal of its own. The worry is that the model's own goal…
Superposition in neural networks is the idea that a model can store more concepts than it has neurons, by overlapping them. It is one reason a single neuron in a large model often responds to several unrelated things,…
Dangerous capability evaluations are tests that ask one narrow question about an AI model: could it do something that would help cause severe harm? Examples are writing working cyber attacks, persuading people against…
Sparse autoencoders are a tool for reading what is going on inside an AI model. A model's neurons each mix many ideas together, so looking at one neuron tells you little. Sparse autoencoders try to pull those mixed…
Weak-to-strong generalization is the question of whether a weak supervisor can train a model that ends up stronger than the supervisor itself. It matters for AI safety because people may one day need to supervise systems…
Inverse reinforcement learning is a way to work out what someone wants by watching what they do. Ordinary reinforcement learning starts with a goal and learns behavior. Inverse reinforcement learning, or IRL, runs the…
Today, much of AI training depends on people judging whether an answer is good. That works while people can tell. Scalable oversight is the research problem of keeping that check working when the AI is better at the task…
If you read about AI safety for long, you meet two terms that sound almost the same: outer alignment and inner alignment. Outer vs inner alignment is not a matter of style. They name two different places where a trained…
A trained neural network does useful things, but nobody wrote down how. Its knowledge sits in millions or billions of numbers. Mechanistic interpretability is the research effort to read those numbers: to explain what a…
Ask what a chess program, a disease-research system and a question-answering assistant have in common, and the obvious answer is very little. Instrumental convergence is the idea that, despite their different goals, all…
Whenever we train an AI system, we give it a number to push up: a score, a reward, a rating. That number is a stand-in for what we actually want. Goodhart's law in AI is the observation that pushing hard on such a…
An AI system can learn its skills perfectly and still learn the wrong goal. That is goal misgeneralization: a trained system keeps its abilities in a new situation but uses them to pursue something other than what it was…
Every machine should have an off switch. For a simple machine that is easy: it has no opinion about being switched off. Corrigibility asks how to keep that property in an AI system that pursues goals, because a system…
Most AI assistants learn their manners from people who rate thousands of answers. Constitutional AI swaps most of that rating for a short written list of principles, and lets an AI model apply the list to its own…
How do you check an answer from a system that knows more than you do? One proposal is to let two AI systems argue about it while a person judges. AI safety via debate turns that idea into a training method: the systems…
Before an AI model is released, someone has to try to make it misbehave. AI red teaming is that job: deliberately attacking a model with tricky, hostile or unusual inputs to find harmful behavior before real users do.…