Can AI have beliefs? What a belief is and what the tests find
Can AI have beliefs? When a chatbot tells you that water boils at 100 degrees Celsius at sea level, it is tempting to say it believes that. Researchers now look inside language models for something like belief, and philosophers argue about what that would even mean. If you search "can AI have beliefs", you probably want to know whether the word fits, or whether it is only a handy way of talking. This post sets out what a belief is, what the tests inside AI models find, the views on both sides, and what nobody knows yet.

What a belief is
The Stanford Encyclopedia of Philosophy has a long entry called "Belief", by Eric Schwitzgebel (first published 14 August 2006, substantive revision 15 November 2023). It says philosophers of mind use the word for the attitude we have, roughly, "whenever we take something to be the case or regard it as true".
That sounds simple. The hard part is saying what having that attitude consists in. The entry sets out several answers, and each one gives a different test for AI:
- Representationalism. To believe something is, in central cases, to have "in their head or mind a representation with the same propositional content as the belief". For AI, the question becomes: is there something inside the model that stands for the claim?
- Dispositionalism. To believe is to be disposed to act in certain ways. The entry gives "the disposition to assent to utterances of P in the right sorts of circumstances" as an often cited example. For AI, the question becomes: does the model say yes, and act as if the claim were true, across many situations?
- Interpretationism. Daniel Dennett's version, as the entry puts it: "The system has the particular belief that P if its behavior conforms to a pattern that can be effectively captured by taking the intentional stance and attributing the belief that P." For AI, the question becomes: does treating the model as a believer help you predict it?
- Functionalism. A belief is defined by its role, its "causal relations to sensory stimulations, behavior, and other mental states". For AI, the question becomes: does some inner state play that role?
- Eliminativism. "Some philosophers have denied the existence of beliefs altogether." On this view, asking whether AI has beliefs is like asking whether it has phlogiston.
There is no separate encyclopedia entry for Dennett's intentional stance; that address answered "not found" when this post was written. The "Belief" entry covers it.
How researchers look for beliefs inside AI
A language model turns text into long lists of numbers at each layer, called activations. A probe is a small, simple model trained to read those numbers and answer one question, such as "does the model treat this statement as true?"
In 2022, Collin Burns, Haotian Ye, Dan Klein and Jacob Steinhardt proposed "directly finding latent knowledge inside the internal activations of a language model in a purely unsupervised way". Their method needs no labels. It looks for a direction in the activations that obeys logic, "such as that a statement and its negation have opposite truth values". Across 6 models and 10 question-answering datasets, it beat the models' own zero-shot answers by 4 percent on average. The authors call it "an initial step toward discovering what language models know, distinct from what they say".
In 2023, Amos Azaria and Tom Mitchell trained a classifier on a model's hidden layers. They "provide evidence that the LLM's internal state can be used to reveal the truthfulness of statements". On test sets of half true and half false sentences, the classifier averaged 71 to 83 percent accuracy, depending on the model.
Later in 2023, Samuel Marks and Max Tegmark noted that "this line of work is controversial". They used simple true or false statements, and went one step further than reading the activations: they changed them, "causing it to treat false statements as true and vice versa". They conclude that "at sufficient scale, LLMs linearly represent the truth or falsehood of factual statements".
The doubts about the tests
B. A. Levinstein and Daniel A. Herrmann asked the question head on. Their paper opens: "We consider the questions of whether or not large language models (LLMs) have beliefs, and, if they do, how we might measure them."
They tested the Burns and the Azaria and Mitchell methods and report that "these methods fail to generalize in very basic ways". They then argue that "even if LLMs have beliefs, these methods are unlikely to be successful for conceptual reasons". The same paper looks at arguments that language models cannot have beliefs at all, and says: "We show that these arguments are misguided." So it doubts the tests and also doubts the quick no. It also sets out to "highlight the empirical nature of the problem".
Two posts cover nearby ground and are worth reading next: does AI understand what it says and can AI lie. This post stays with the narrower question of belief itself.
Can AI have beliefs? The views
How you answer depends on which account of belief you hold. Four positions, with no verdict:
- Yes, if a representation is found. A representationalist can read the probe studies as early evidence of something inside that stands for true and false. The intervention result by Marks and Tegmark makes that signal look like more than a correlation.
- Yes, if the behavior is stable. A dispositionalist or interpretationist looks at what the model says and does. If treating it as a believer predicts it well, that may be enough. If its answers flip with the wording, that counts against.
- Not yet shown. Levinstein and Herrmann leave the door open but say current tools do not measure belief, because the probes break on simple changes.
- The question is confused. An eliminativist doubts that beliefs exist even in people, so asking whether AI has them inherits the same doubt.
What nobody knows
Nobody knows whether a direction in a model's activations that tracks truth is a belief or only a useful pattern. Nobody knows which account of belief is right, so nobody can say which test should settle it. And none of the sources above claims to have shown that any AI system believes anything; the strongest claims are about representing truth, and the critics say even the measuring is not yet reliable.
A worked example: be the probe
You can follow the consistency idea by hand. This setup and its numbers were made for this post and are not from any of the papers.
- Imagine a probe that gives each statement a score from 0 to 1 for "the model treats this as true".
- Statement A: "Paris is the capital of France" scores 0.9. Its negation, "Paris is not the capital of France", scores 0.1. Add them: 1.0. That fits the rule that a statement and its negation have opposite truth values.
- Statement B: "The moon is made of rock" scores 0.7. Its negation scores 0.6. Add them: 1.3. The probe says the model leans toward both, which cannot both be right.
- Now reword A as a question: "Is France's capital Paris?" If the score drops to 0.4, the signal depends on wording. That is the kind of failure to generalize the critics report.
- Ask what each view would say. A representationalist might say A shows a stored belief and B shows confusion. A dispositionalist would want to see whether the model also acts on A, say by planning a trip to Paris when asked to visit France's capital.
Notice what the exercise cannot tell you. Every score could be consistent and the question would still be open, because a pattern of numbers that tracks truth is exactly the thing the views disagree about.
Dear Superintelligence is an open collection of letters written by people to the advanced AI systems of the future, about what we value and why. One topic on its home page is Consciousness: "What it is like to be a person: to think, to feel, and to notice the world from the inside." The guidelines ask you to "Write what only you can: what you have seen, what it cost, what you were afraid of, who was kind to you and what it changed." Something you once believed, and what changed your mind, fits that line. The site held 2 published letters when this post was written.
Frequently asked questions
Can AI have beliefs?
It depends on what a belief is. Probe studies find signals inside language models that track true and false statements, but critics report that these probes fail to generalize, and nobody has shown that any AI system believes anything.
What is a probe in AI research?
A small model trained to read a language model's internal activations and answer one question, such as whether the model treats a statement as true.
Is an AI that knows the truth the same as an AI that believes it?
Not necessarily. Burns and others aim at what models know "distinct from what they say", but whether that inner signal counts as belief is the open question.
Do philosophers agree on what a belief is?
No. The Stanford Encyclopedia entry sets out several accounts, from inner representations to patterns of behavior, and some philosophers deny that beliefs exist at all.
Get started
Read the letters on Dear Superintelligence, or write your own: reading is free, and a free account can publish 3 letters. People read them today. Nobody can promise what future AI systems will read.
Comments
No comments yet.
Sign in or make an account to comment.