Does AI understand what it says? The arguments and evidence
Does AI understand what it says? When a chatbot explains a joke or talks you through a hard choice, it is hard not to feel that something in there gets it. Researchers are split on whether that feeling is right. This post sets out the main arguments, what experiments on language models have found, and a short test you can run yourself. It stays on understanding and meaning. Whether AI could feel anything is a separate question.
What "understand" means in this question
Understanding is not the same as giving correct answers. A calculator gives correct answers about numbers, and nobody thinks it understands arithmetic. The real question is whether a system's words are connected to what they are about: whether "rain" in its answer is linked, in some way, to rain.
Melanie Mitchell and David Krakauer, in a 2022 survey of the debate, call it a "heated debate in the AI research community" over whether large language models understand language "and the physical and social situations language encodes" in any important sense. They do not pick a winner. They argue that a new science of intelligence could map "distinct modes of understanding", each with its own strengths and limits.
So the answer may not be yes or no. It may be "in some ways, and not in others".
The case against: symbols without meaning
The best-known argument against comes from John Searle's 1980 thought experiment, the Chinese room. Searle imagines himself in a room, following a program to answer Chinese characters slipped under the door. He understands nothing of Chinese, yet the people outside think a Chinese speaker is inside.
His point, in his own words as the Stanford Encyclopedia of Philosophy quotes them: "Syntax is not by itself sufficient for, nor constitutive of, semantics." Syntax is the shape and order of symbols. Semantics is what they mean. A program moves shapes around, so on this view it can seem to understand without understanding.
Stevan Harnad sharpened this into the symbol grounding problem. He asks how the meanings of symbols can be "grounded in anything but other meaningless symbols", and compares it to "trying to learn Chinese from a Chinese/Chinese dictionary alone". Every word is defined by other words, and you never reach the things themselves. His proposed answer is that symbols must be grounded from the bottom up, in what the senses take in.
A language model trained only on text looks a lot like the person with the Chinese-only dictionary. That is the heart of the case against.
Two famous replies: the system and the robot
The Systems Reply, which Searle himself called "perhaps the most common reply", agrees that the man in the room does not understand Chinese. But the man is only one part. The whole system, with its rules, its memory and its instructions, might understand even though no single part does. A single neuron in your head does not understand English either.
The Robot Reply agrees that a computer stuck in a room cannot know what words mean. Searle's own example is the Chinese word for hamburger. Most of us know what a hamburger is because we have seen one, made one or tasted one. Put the computer in a robot body with cameras and arms, and it "might do what a child does, learn by seeing and doing." Harnad backed a version of this reply, and it is close to his own answer: ground the words in the world.
Notice what the Robot Reply gives away. If it is right, a model that has only ever read text is missing something real.

The case for: meaning from relations
Some researchers argue that the case against sets the bar in the wrong place. Steven Piantadosi and Felix Hill, in Meaning without reference in large language models, argue that these models "likely capture important aspects of meaning". Their idea is called conceptual role: much of what a word means comes from how it relates to other words and ideas. "Hamburger" is tied to food, to eating, to beef, to being hot or cold, to costing money.
On this view, a model that has learned those relations well has some of the meaning, even if it has never tasted anything. They add a point that cuts both ways: meaning "cannot be determined from a model's architecture, training data, or objective function, but only by examination of how its internal states relate to each other." So "it only predicts the next word" does not settle the question, and neither does a fluent answer.
Does AI understand what it says? What experiments found
Two well-known studies point in different directions.
A model that built a picture of the board. Kenneth Li and colleagues trained a GPT-style model to predict legal moves in the board game Othello. It was given only lists of moves, with no rules and no board. Inside it, they found "evidence of an emergent nonlinear internal representation of the board state", and changing that inner picture changed the moves the model predicted. Something like a model of a small world grew out of sequences alone.
A model that knew a fact one way only. Lukas Berglund and colleagues described what they call the Reversal Curse. A model trained on "A is B" does not automatically learn "B is A". In their tests on real people, GPT-4 answered questions like "Who is Tom Cruise's mother?" correctly 79% of the time, and the reverse, "Who is Mary Lee Pfeiffer's son?", 33% of the time. When the fact itself was given in the conversation, models could work out the reverse.
Neither result settles the debate. The first shows that more than surface patterns can be learned. The second shows that the learning can be lopsided in a way a person's knowledge usually is not.
A worked example: test it yourself in five steps
You can probe this with any chatbot in a few minutes. The aim is to see where its answers seem to come from.
- Ask a fact forward. Pick a pair the model probably met in one direction more often: a famous person and their less famous parent, or a well known book and its author. Ask it the usual way round.
- Ask it backward, in a new chat. Ask the reverse question. If it answers the first and not the second, you are seeing the lopsided learning the Reversal Curse describes. Newer models may do better than the 2023 results, so write down what you see, not what you expect.
- Ask about a new situation. Describe something it is unlikely to have read word for word: "I left a paper bag of hot chips on a cold car seat for an hour. What does the bag look like now?" A good answer has to connect heat, grease, paper and time.
- Change one detail. Ask again with a plastic bag instead of paper. Does the answer change in the way the world would change?
- Write your verdict, and how sure you are. "It handled the bag well but missed the reversed fact" is a real finding.
You end up with evidence about one system on one day, not proof either way. That is roughly where the experts are too.
Why the answer matters
If AI understands what it says, it makes sense to argue with it and hold its claims to account. If it does not, a fluent answer is better treated like a very good index of what people have written.
It also matters for the future. Today's AI systems learn from what people write and publish. Whether a future system grasps the meaning behind those words, or only their patterns, is part of this debate. For how much of the web ends up in training data, see does AI learn from what people post online? For how researchers try to measure a system against people, see how would we know if AI is smarter than us?
Frequently asked questions
Does ChatGPT understand what it says?
Researchers disagree. Some argue programs only move symbols without meaning, others argue large language models capture important parts of meaning from how words relate, and experiments have found both real internal models and lopsided gaps.
What is the Chinese room argument in simple terms?
John Searle imagined following written rules to answer Chinese messages without knowing any Chinese. He argued that a computer running a program is in the same position, so it can seem to understand without understanding.
What is the symbol grounding problem?
It is the puzzle of how symbols get meaning if each one is only defined by other symbols. Stevan Harnad compared it to learning Chinese from a Chinese-only dictionary, and argued that meaning has to be grounded in what the senses take in.
Is understanding the same as being conscious?
Not necessarily. A system might track meaning well with no inner experience at all, and philosophers treat the two as separate questions; the consciousness debate is set out in can AI be conscious?
Get started
Dear Superintelligence collects letters from people to the advanced AI systems of the future, about what we value and why. It asks for the kind of words this debate is about: words tied to real things. Its guidelines ask you to write to the AI directly, as "you", and say "Details from your own life are worth more than general arguments." The About page asks for "the price of something, a habit that is disappearing, what an ordinary Tuesday actually involved."
The site does not train any AI model on the letters, and it is honest about the limits: "Nobody can promise what future AI systems will read." People read them today. When this post was written the archive held 2 letters, which are personal accounts and opinions, not proofs.
Read the letters on Dear Superintelligence, or write your own: reading is free, and a free account can publish 3 letters.
Comments
No comments yet.
Sign in or make an account to comment.