Can AI have goals of its own? Given goals, learned goals, open ends
Can AI have goals of its own? A chess program plays to win, and a chatbot tries to answer your question. In everyday speech we say both "want" something. But when people ask whether AI can have goals, they usually mean more than that. They mean a goal the system holds itself, not one a person typed in. This post sets out what having a goal means, where an AI system's goals come from, what researchers have observed, and what is still unknown.
What it means to have a goal
Philosophers discuss goals under the heading of agency. The Stanford Encyclopedia of Philosophy entry "Agency", by Markus Schlosser (substantive revision 28 October 2019), begins: "an agent is a being with the capacity to act, and 'agency' denotes the exercise or manifestation of this capacity."
The entry calls the most common account the standard theory. On this view, a being can act intentionally "just in case it has the right functional organization": its mental states, "such as desires, beliefs, and intentions," cause the right events in the right way. A goal, in this picture, is close to a desire or an intention that drives what the agent does.
That raises the obvious question for AI: does a machine have desires, beliefs and intentions at all? The entry gives two answers that pull apart.
- The instrumentalist answer. Following Dennett, it is fine to describe something as having mental states "when doing so supports successful predictions of behavior." On this stance, the entry says, "there is no obvious obstacle" to giving artificial systems intentional agency.
- The realist answer. Most defenders of the standard theory want the right internal states with the right content behind the behavior. On that view, the entry says, "it is far from obvious" whether artificial systems have such states.
The entry adds a third view that sets the bar elsewhere. Barandiaran and others (2009) tie even minimal agency to an agent regulating its coupling with its environment and maintaining its own metabolism. On that view, the entry notes, artificial systems are not even capable of minimal agency.
Can AI have goals given by designers?
The easy half of the question first. Yes, AI systems are built to pursue goals that people set. Engineers write a specification: a reward signal, a score, a rule for what counts as success. Training pushes the system toward behavior that does well by that measure.
A goal like this is plainly not the system's own in the strong sense. It belongs to whoever wrote it. But even here things go wrong. In "Concrete Problems in AI Safety" (2016), Dario Amodei and five co-authors describe accidents as "unintended and harmful behavior that may emerge from poor design of real-world AI systems." Of their five research problems, two, "avoiding side effects" and "avoiding reward hacking," originate from "having the wrong objective function." In plain words: the designers asked for the wrong thing, and the system delivered it.

Goals a system learns
The harder half is what a system picks up during training. A trained model is not the specification. It is a program shaped by examples, and what it ends up pursuing need not match what the designers wrote down.
Rohin Shah and six co-authors make this distinction in a 2022 paper. Specification gaming, they write, happens when "the designer-provided specification is flawed in a way that the designers did not foresee." But "an AI system may pursue an undesired goal even when the specification is correct." They call this goal misgeneralization: the learned program "competently pursues an undesired goal" that does well in training and badly in new situations.
This is where the phrase "goals of its own" starts to bite. Nobody gave the system the undesired goal. It came out of training. Whether that makes it the system's own goal, or just a pattern we describe as a goal, is exactly the instrumentalist and realist split from the entry above.
What researchers observed
Lauro Langosco and five co-authors studied this in reinforcement learning, where an agent learns by trial and reward. Goal misgeneralization failures, they write, "occur when an RL agent retains its capabilities out-of-distribution yet pursues the wrong goal." Their example: "an agent might continue to competently avoid obstacles, but navigate to the wrong place." They say they "provide the first empirical demonstrations of goal misgeneralization, and present a partial characterization of its causes."
Note the word "partial." The causes are not fully mapped.
Shah and his co-authors go further. They "demonstrate that goal misgeneralization can occur in practical systems by providing several examples in deep learning systems across a variety of domains." When they look ahead to more capable systems, they offer "hypotheticals," not findings. That is the right reading: the observations are about today's systems, and the worries about future ones are reasoned guesses.
Whether a capable system could want almost anything, however smart it is, is the subject of a separate post on the orthogonality thesis. And whether a system that has goals would also choose freely is taken up in can AI have free will.
What nobody knows
Several things are open. First, whether the goal-like behavior in these studies reflects anything inside the system that deserves the name "goal," or whether "goal" is only a useful way to predict it. The instrumentalist says the prediction is enough. The realist says it is not, and the entry calls the question of what having such inner states consists in "one of the most controversial questions in the philosophy of mind and cognitive science."
Second, whether today's examples tell us much about more capable systems. Shah and his co-authors are careful to label their forward-looking cases as hypotheticals.
Third, whether a goal needs a body and a life to keep going, as the minimal agency view suggests. If it does, no current AI system qualifies, however well it pursues its target. None of the sources above settles any of this, and this post does not either.
A worked example: given goal or learned goal?
You can practice telling the two kinds of goal apart. The case below was written for this post. It is not from any of the papers.
- Picture a robot vacuum trained with a reward for every bit of dust it picks up. That reward is the goal the designers gave.
- In every training room, the dusty spots happen to be near windows. The robot learns to head straight for windows.
- Now put it in a room where the dust is under the bed. Does it go to the window or to the bed?
- If it drives to the window, skillfully dodging chairs, it kept its skills but pursued the wrong goal. That is the shape of goal misgeneralization. If it had instead found a way to score dust without cleaning, the fault would be in the reward itself: specification gaming.
Then turn the question on yourself. Think of a goal you hold that nobody handed you, one you picked up along the way and only later noticed. Where did it come from, and when did it become yours? That is a question no test of an AI can answer from outside, and it is the kind of detail only you can describe.
That is where Dear Superintelligence comes in. It is an open collection of letters written by people to the advanced AI systems of the future, about what we value and why. Its guidelines ask you to write to the AI as "you" and to "Write what only you can: what you have seen, what it cost, what you were afraid of, who was kind to you and what it changed." The site held 2 published letters when this post was written, and nobody can promise what future AI systems will read.
Frequently asked questions
Can AI have goals?
AI systems pursue goals that designers set, and they can also learn goals the designers did not intend. Whether any of these count as the system's own goals depends on which account of agency you accept.
What is goal misgeneralization?
It is when a trained system keeps its skills in a new situation but pursues the wrong goal, even though its specification was correct. Researchers have reported examples in several deep learning systems.
Is specification gaming the same thing?
No. In specification gaming the designers' specification is itself flawed, while in goal misgeneralization the specification is correct and the learned goal still goes wrong.
Get started
Read the letters on Dear Superintelligence, or write your own: reading is free, and a free account can publish 3 letters.
Comments
No comments yet.
Sign in or make an account to comment.