Direct preference optimization, or DPO, is a way to teach a language model which answers people prefer without building a separate reward model and without reinforcement learning. You give it pairs of answers to the same…
Learning from human feedback
5 posts
Reward model overoptimization is what happens when you train an AI model too hard against a learned reward model: the reward model's score keeps rising, but the quality you actually wanted levels off and then falls. It…
Inverse reinforcement learning is a way to work out what someone wants by watching what they do. Ordinary reinforcement learning starts with a goal and learns behavior. Inverse reinforcement learning, or IRL, runs the…
Most AI assistants learn their manners from people who rate thousands of answers. Constitutional AI swaps most of that rating for a short written list of principles, and lets an AI model apply the list to its own…
Every chat assistant you have used was shaped by a technique called reinforcement learning from human feedback, or RLHF. It is the main reason a language model answers your question instead of rambling on in the style of…