Eliciting latent knowledge, usually shortened to ELK, is the problem of getting an AI to tell you what it actually knows, rather than what it predicts you would believe. That sounds like a small difference. The…
Scalable oversight
4 posts
Eliciting latent knowledge explained: the SmartVault problem
Weak-to-strong generalization: can weak AI supervise strong AI?
Weak-to-strong generalization is the question of whether a weak supervisor can train a model that ends up stronger than the supervisor itself. It matters for AI safety because people may one day need to supervise systems…
Scalable oversight: how humans could supervise smarter AI
Today, much of AI training depends on people judging whether an answer is good. That works while people can tell. Scalable oversight is the research problem of keeping that check working when the AI is better at the task…
AI safety via debate: how two AIs arguing could help
How do you check an answer from a system that knows more than you do? One proposal is to let two AI systems argue about it while a person judges. AI safety via debate turns that idea into a training method: the systems…