Untrusted monitoring is a safety technique from the field of AI control: you use one copy of a powerful AI model to check the work of another copy, and assume either copy might be trying to trick you. It sounds…
Alignment theory
4 posts
Untrusted monitoring: how AI control keeps models in check
Embedded agency explained: an AI inside its own world
Embedded agency is the problem of how an agent should reason and act when it is part of the world it is trying to change. Most theories of rational choice quietly assume the agent sits outside its world, like a player…
Instrumental convergence explained simply
Ask what a chess program, a disease-research system and a question-answering assistant have in common, and the obvious answer is very little. Instrumental convergence is the idea that, despite their different goals, all…
Corrigibility: why an AI that accepts correction is hard to build
Every machine should have an off switch. For a simple machine that is easy: it has no opinion about being switched off. Corrigibility asks how to keep that property in an AI system that pursues goals, because a system…