Natural Abstraction Hypothesis Explained: Will AI Think Like Us?
If you have read alignment forums for long, you have probably seen the phrase "natural abstractions" and wondered what it means. This post explains the natural abstraction hypothesis in plain terms. It covers what John Wentworth claims, why the idea comes from physics, why alignment researchers care, and where the evidence and the objections stand today. By the end you should be able to state the hypothesis in one sentence and say what would count as evidence for or against it.
What is the natural abstraction hypothesis?
The hypothesis is a bet about the world, not about AI. It says our universe has a set of "natural" high-level concepts that almost any good learner will discover. Trees, rocks, temperature, people, and cars are all examples. On this view, these concepts are not arbitrary human inventions. They are summaries that the structure of the world almost forces on you if you want to predict it well.
If that is true, a large AI model trained on data from our world should end up with internal concepts that look a lot like ours. That matters because some alignment plans depend on finding and reading the concepts inside a model. If the model thinks in alien terms, those plans get much harder.
The three claims: abstractability, human-compatibility, convergence
In his 2021 post Testing the Natural Abstraction Hypothesis: Project Intro, Wentworth split the hypothesis into three claims. Each can be true or false on its own, so it helps to keep them apart.

- Abstractability. The physical world abstracts well. For most systems, the information that matters to things far away is much smaller than the full detail of the system. You can throw away almost everything and still predict well at a distance.
- Human-compatibility. The small summaries that survive this compression are, roughly, the concepts humans already use. Our word "tree" points at something real and compact, not at a cultural accident.
- Convergence. A wide variety of cognitive architectures, which could mean human brains or neural networks, learn and use approximately the same summaries when they model the same world.
Notice that the first claim is about physics, the second about people, and the third about learning systems in general. Wentworth calls the first two empirical claims, to be tested in the real world, and the third a more mathematical one. You could accept abstractability and still doubt convergence.
Information at a distance: how abstractions emerge from physics
The core intuition is information at a distance. Wentworth's own example is a box of gas: a huge number of particles, each with a position and a velocity. Almost none of that detail lasts. Collisions scramble it, and chaos amplifies even tiny uncertainty about the starting state. What survives is a handful of numbers: the energy, the number of particles and the volume of the box. That, he notes, is exactly the abstraction behind the ideal gas law.
Those few numbers are the abstraction. They are not chosen because humans find them convenient. They are what is left after the world itself has washed out the noise: in the post's words, everything is lost "except for information which is perfectly conserved."
Later work by Wentworth and David Lorell, Natural Latents (2025), takes on the convergence claim with math. It asks: if two agents predict the same world equally well but use different internal variables, when can one agent's variables be translated into the other's? It gives conditions, the natural latent conditions, under which that translation is always possible, and shows the result holds even when the conditions are only roughly met.
A worked example you can try
Run this thought experiment with the concept "peach."
- List the low-level facts about one peach: atom positions, the exact shape of each cell, the bruise pattern.
- Ask which facts matter to things far away. A person across the room cares about color, size, ripeness, and location. The cell layout almost never matters at a distance.
- Ask which facts are redundant. Color, smell, and texture all carry information about ripeness. Ripeness is copied across many channels, so it survives noise.
- Now imagine an alien learner, or a vision model, trained only to predict what happens to fruit. Ask whether it would need a "ripeness" variable to do well.
If your answer to step 4 is "almost certainly yes," you have just felt the pull of the convergence claim. If you can think of a good predictor that never forms a ripeness concept, you have found a possible counterexample. That is the test the hypothesis invites.
Why natural abstractions matter for AI alignment
Alignment needs a way to point an AI at what we mean. Suppose you want a system to care about "human wellbeing" or to report "whether the diamond is still in the vault." Somewhere inside the model there has to be a concept that matches yours. Wentworth's post makes the point directly: human values are about things like trees, cars and other humans, not about low-level states of the world.
If the natural abstraction hypothesis holds, his post says, it "would dramatically simplify" alignment. The model's concepts should line up with ours, at least for everyday physical things. You could then search for them, read them, and maybe steer the model through them. If the hypothesis fails, the model's internal concepts could be carved up in ways no human would recognize, and asking "what does it believe about X?" might have no clean answer.
This links directly to the problem of eliciting latent knowledge. In the SmartVault setup, you want an AI to tell you what it actually knows about the diamond, not what a human would believe from the camera feed. Natural abstractions would make the honest concept easier to locate, because "diamond is in the vault" would be a compact, convergent variable inside the model.
Evidence so far: interpretability, linear probes and shared representations
The hypothesis makes a testable prediction: trained models should contain human-like concepts in readable form. Several lines of interpretability work bear on this.
- Linear probes. Researchers train a simple linear classifier on a model's internal activations to see if a concept can be read off. If you want the method and its limits, see linear probes explained.
- Features. Sparse autoencoders try to pull readable features out of a model, so that features in different models can be compared.
- Converging representations. A 2024 paper, The Platonic Representation Hypothesis, argues that representations in AI models are converging, and reports that as vision and language models get larger, they measure the distance between data points in more and more alike ways.
All of this is consistent with convergence. None of it proves the hypothesis. Probes can find a concept that the model does not actually use, and "similar" representations can still differ in ways that matter for safety. Treat the evidence as suggestive, not settled.
Criticisms and open problems with the hypothesis
The hypothesis faces hard questions, and fair treatment means stating them plainly.
- Abstractions depend on goals. What counts as "far away" or "relevant" depends on what you care about. A chemist and a chef carve up the same kitchen differently. Is there one natural set or many?
- Human concepts may vary. If different groups of people carve up social and abstract ideas differently, which carving is the "human" one?
- Values may not be natural. Even if "tree" is natural, concepts like "fairness" or "wellbeing" might not be compact summaries of physics. Those are the ones alignment most needs.
- Superhuman concepts. A stronger system might find abstractions humans have never formed. Convergence on everyday objects says little about what happens beyond human understanding.
- Approximate is not exact. The claim is "approximately the same" abstractions. Small mismatches could still be enough for an AI to satisfy the letter of a goal and miss its intent.
Wentworth himself frames it as "an empirical claim" that has to be tested, which is why his project was named Testing the Natural Abstraction Hypothesis. How far it extends is an open research question.
Frequently asked questions
Who proposed the natural abstraction hypothesis?
John Wentworth set it out on the Alignment Forum, in an April 2021 project post, as three claims to be tested. Later work with David Lorell developed the math of natural latents.
Is the natural abstraction hypothesis true?
Nobody knows yet. Some results fit the convergence claim, but there is no proof, and abstract or value-laden concepts are the hardest case.
How does it relate to eliciting latent knowledge?
Eliciting latent knowledge asks how to get a model to report what it really knows. If natural abstractions hold, the relevant concept should exist inside the model in a compact, findable form, which makes that problem more tractable.
Would natural abstractions make alignment easy?
No. It would make one step easier: finding the model's concepts. You would still need to make the model act on the right concepts, stay honest, and handle goals that are not natural abstractions.
Get started: learn AI alignment theory step by step
Natural abstractions sit where several bigger topics meet: how goals are specified, how we read a model's internals, and how we oversee systems smarter than us. Learn AI Alignment Theory covers that ground in short lessons of about 8 minutes. The Advanced level includes courses on Agent Foundations, Interpretability, and Scalable Oversight, and Scalable Oversight has a lesson on Latent Knowledge. Each lesson separates what is known from what is still open and lists its sources. Where researchers disagree, debate cards set out each serious position with no verdict.
If you are new, start at the Basic level with What Is Alignment? and Specifying Goals, then work up. Questions from lessons you finish come back on a spaced schedule so the ideas stick. You sign in with Google or an emailed code. You can read more on the about page.
Start Learn AI Alignment Theory: short hands-on lessons from first principles to open research, with every side of every debate and the sources to read next.
Comments
No comments yet.
Sign in or make an account to comment.