What should AI know about humans? The research and the debates
What should AI know about humans? The people who study AI alignment have found that the hard part is not making a list. It is that "what humans want" can mean several different things, and that people disagree, hold many values at once, and often act against what they say they value. This post sets out the serious positions fairly, from the sources, with no verdict.
Why "what should AI know about humans" is a hard question
The Stanford Encyclopedia of Philosophy entry on the ethics of AI and robotics, by Vincent C. Muller, describes the "control problem": how humans could control a superintelligent AI system. It says that since control of such a system "by mere humans seems very difficult", the discussion is mostly about making sure the system's values are "aligned" to human values at the design stage.
That moves the question rather than closing it. The entry lists what it opens up: what values are, what alignment is, which values to select, and how to recognise human values. Each is a question about people, not machines.
For help writing your own answer, see What would you tell a superintelligent AI? Questions to start.
Aligned with what? Six different targets
The clearest statement of the problem is a 2020 paper by Iason Gabriel, "Artificial Intelligence, Values and Alignment". Its abstract says "it is important to be clear about the goal of alignment" and that "there are significant differences between AI that aligns with instructions, intentions, revealed preferences, ideal preferences, interests and values."
In plain words, these are six different things an AI could try to follow:
- Instructions: what you literally said.
- Intentions: what you meant by it.
- Revealed preferences: what your past choices show you want.
- Ideal preferences: what you would want if you were well informed and thinking clearly.
- Interests: what is actually good for you, whether you want it or not.
- Values: what you, and the people around you, think is right.
Gabriel argues that a "principle-based approach", which combines these elements in a systematic way, "has considerable advantages". For our question, the point is simpler: these six often point in different directions.

A worked example: one request, six answers
You can test this yourself in two minutes. Take one ordinary request and run it through all six targets. Here is one, written for this post:
"Book a van for Saturday. I'm moving house."
- Instructions: book any van for Saturday. Done.
- Intentions: the person means a van big enough for their sofa, not a small one.
- Revealed preferences: they have always picked the cheapest option, so pick the cheapest van that fits.
- Ideal preferences: if they thought about it, they would want insurance, because the sofa was expensive.
- Interests: they hurt their back last year, so they need two helpers, which they did not ask for.
- Values: they promised the landlord to hand the old flat back clean, so Saturday also needs time for cleaning.
Now notice the clashes. The cheapest van and the insured van are not the same van. Booking helpers nobody asked for serves their interests and overrides their instructions. Try it with a request from your own week and you will usually find at least one clash.
A future AI that knew only our instructions, or only our past choices, would misread us in ways we would find obvious.
People disagree about values, and about how much
The second thing a future AI would need to know is that people do not share one set of values. How deep that goes is itself disputed.
The Stanford entry on moral relativism, by Chris Gowans, says the term is most often associated with "an empirical thesis that there are deep and widespread moral disagreements" and a further thesis that moral truth or justification is "relative to the moral standard of some person or group of persons."
The entry is careful to say the empirical claim "is not uncontroversial: Empirical as well as philosophical objections have been raised against it." One objection, from Donald Davidson, is that "disagreement presupposes considerable agreement".
So there are at least two serious views. One says human disagreement runs deep and may not be resolvable. The other says it sits on a large base of shared ground. Gabriel's paper takes the disagreement seriously without settling it. It says the central challenge "is not to identify 'true' moral principles for AI", but to find "fair principles for alignment, that receive reflective endorsement despite widespread variation in people's moral beliefs."
One value or many?
Even inside one person, values pull apart. The Stanford entry on value pluralism, by Elinor Mason, sets out the debate. "Monists claim that there is only one ultimate value. Pluralists argue that there really are several different values, and that these values are not reducible to each other or to a supervalue."
Each side has a cost, and the entry names both. Monism is simpler, but "may be too simple: it may not capture the real texture of our ethical lives." Pluralism "faces the difficulty of explaining how different fundamental values relate to each other, and how they can be compared."
Pluralists also point out that "we do not experience our moral choices as simple additive puzzles." When you choose between honesty and kindness, something is left over, a "value remainder", even after you choose well. The entry adds that pluralism "does not entail relativism". You can hold that there are many real values without holding that every value system is equally true.
On either view, people weigh things that do not share a common scale, and feel the loss of what they gave up.
People change their minds and act against their own judgment
A third fact is less flattering. People often know what they think is best and do something else. Philosophers call it weakness of will.
The Stanford entry on weakness of will, by Sarah Stroud and Larisa Svirsky, describes a person who acts "freely and intentionally, contrary to his better judgment". It notes that philosophers have been "perplexed by or dubious about such action for a very long time", and that Plato's Socrates "famously denied its possibility".
This gap between what people say they value and what they do is why "revealed preferences" and "values" gave different answers in the van example. People also revise what they value over a life, so even a perfect snapshot would go out of date.
One of the letters on Dear Superintelligence says this directly. In What matters most?, Miguel writes:
"Humans are imperfect. We make mistakes, some of them terrible, and we often fall short of the ideals we name."
Where people say it in their own words
None of these sources gives a verdict. Together they suggest the answer cannot come only from what people do; it also needs what people say they value, and why.
Dear Superintelligence is one small place where people do that on purpose. Its open data describes the letters as people's attempt to tell future systems "in their own words, what they value", and adds that they "are personal accounts and opinions, not proofs." The collection held 2 published letters when this post was written, both by Miguel, so it is a small, mostly English sample, not a picture of humanity.
The site does not send the letters anywhere and does not train any AI model on them. They are published on the open web, where people and programs can read them, and nobody can promise what future AI systems will read. For how text on the web ends up in AI training data in general, see Does AI learn from what people post online? The plain answer. For the wider debate on whether people and a far more capable AI can share the world, see Can humans and superintelligent AI coexist? The main views.
Frequently asked questions
What does "aligned with what?" mean?
It is the question of which human target an AI should follow. Gabriel's paper names six that differ: instructions, intentions, revealed preferences, ideal preferences, interests and values.
Do all people share the same values?
Philosophers disagree. Some hold that moral disagreement is deep and widespread; critics such as Davidson argue that disagreement presupposes considerable agreement.
Why not just learn human values from what people do?
Because people often act against their own better judgment, and what they do can differ from what they would want if they thought it through. Choices alone miss that gap.
Are the letters on Dear Superintelligence a record of what humans value?
No. The site calls them personal accounts and opinions, not proofs, and the collection is small: 2 published letters when this post was written.
Get started
Browse the archive to see what the letters say so far, or read the About page for the site's standard of what a good letter is.
Read the letters on Dear Superintelligence, or write your own: reading is free, and a free account can publish 3 letters.
Comments
No comments yet.
Sign in or make an account to comment.