Plenty of people want to understand AI safety, whether to change careers, to work on policy, or simply to follow the arguments properly. Most start with a reading list, open the first paper, and stall somewhere around…
AI alignment
35 posts
AI alignment is the work of making AI systems do what we intend. That sounds simple, and for a spreadsheet macro it is. For systems that learn their behaviour from data and rewards rather than following rules someone…
Every chat assistant you have used was shaped by a technique called reinforcement learning from human feedback, or RLHF. It is the main reason a language model answers your question instead of rambling on in the style of…
Few ideas in AI safety are argued about as hotly as deceptive alignment, also called scheming: the possibility that an AI system could behave well during training because it understands it is being trained, while…
Ask a cleaning robot to minimise the mess it can see, and the easiest solution may be to cover its camera. Nothing has gone wrong with the robot. It found the cheapest way to score well on the objective it was given.…