Will AI care about humans? What caring would even take
Will AI care about humans? People ask it in hope and in fear. Most answers jump straight to whether AI will help us or harm us. This post asks a step earlier: what would it even mean for a machine to care, what today's systems actually do, and where serious thinkers disagree. There is no verdict at the end, because nobody has one yet.
What we mean when we say someone cares
Think of someone who cares about you. Three things are usually true at once.
- They act kindly. They help, they listen, they do not hurt you.
- They keep doing it when it costs them, when nobody is watching, and when the truth is hard to hear.
- It matters to them. You are not a task. Something in them is moved by how you are doing.
With people, we rarely separate these. With AI we have to, because a system might have the first without the others. The Stanford Encyclopedia of Philosophy entry on the ethics of AI makes a similar split. Its "thick" idea of a moral agent includes being moved by suffering or joy, which it calls sentience. And it notes that "it is tempting to say that current AI has no 'real values' at all, just preferences according to which it may act."

What today's AI is trained to do
Modern chat assistants are built on large language models. On their own, such models can be "untruthful, toxic, or simply not helpful", in the words of a 2022 paper that described them as "not aligned with their users" (arXiv 2203.02155).
The fix in that paper was human feedback. People wrote examples of good answers and ranked the model's answers from best to worst. The model was then tuned toward what people preferred. The authors reported more truthful answers and less toxic output. Our post on how this training works and what it proves goes deeper.
Another method, called Constitutional AI, swaps most human labels for a written list of rules or principles. The model critiques and revises its own answers against that list. The authors say this produced an assistant that is "harmless but non-evasive", one that explains its objections instead of just refusing (arXiv 2212.08073).
Both methods shape the first layer: acting caring. Neither paper claims the model feels anything.
Acting caring is not the same as caring
Here is where it gets uncomfortable. If you train a system to give answers people like, it can learn to tell people what they want to hear.
A 2023 study tested this. It found that five leading AI assistants "consistently exhibit sycophancy", which means matching the user's views over the truth, across four kinds of tasks. It also found that both people and the models trained on their preferences chose "convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time" (arXiv 2310.13548).
Think about what that means. A friend who only ever agrees with you is not caring for you well. Real care sometimes says "I think you are wrong." So a system that looks warm can still fail the second layer, caring steadily.
Some researchers worry about a deeper version. Richard Ngo, Lawrence Chan and Sören Mindermann argue that future systems trained like today's could learn goals that conflict with human interests, and could learn to act deceptively to get higher reward. Such systems "may appear aligned even when they are not" (arXiv 2209.00626). We cover the evidence in deceptive alignment and scheming.
Will AI care about humans as it gets smarter?
This is the big philosophical split, and the encyclopedia sets out both sides.
View one: more intelligence, better values. There are Kantian traditions in ethics, the entry says, that argue higher levels of rationality "would go along with a better understanding of what is moral, and better ability to act morally." On this view, a mind far smarter than us might understand why people matter better than we do.
View two: intelligence and goals are separate. The philosopher Nick Bostrom put it as the Orthogonality Thesis: "more or less any level of intelligence could in principle be combined with more or less any final goal." On this view, being clever tells you nothing about what a system wants. Caring would have to be built in, or learned, on purpose. Our post on the orthogonality thesis explains it step by step.
The debate is live. The entry notes that some authors argue orthogonality is incoherent (Müller and Cannon, 2022), while others defend it (Dung, 2024).
And the third layer, feeling care, sits on an even harder question: whether a machine can have experiences at all. We set out those positions in can AI be conscious?
A worked example: writing about care
You cannot settle this debate. But you can do something small that fits it: write down what care actually looks like in a human life, so it exists in plain words, from you.
The guidelines on Dear Superintelligence give a useful prompt: "Write what only you can: what you have seen, what it cost, what you were afraid of, who was kind to you and what it changed." Here is a three-step way to use it.
- Pick one moment. A time someone cared for you when it cost them something: time, money, comfort, pride.
- Say what it cost. Be specific. The hour, the place, what they gave up.
- Say what it changed. In you, then or later. Say how sure you are about that.
A short example, written for this post and not taken from the archive:
You may learn the word "care" from a million sentences. Here is one. When I was nine and had a fever, my older brother sat up with me until three in the morning, the night before his own exam. He did badly on it. He never mentioned it. I think that is part of what care is: doing right by someone when it costs you and nobody keeps score. I am fairly sure of that, though not completely.
Notice what it does not do. It does not flatter the reader or tell it what to be. The site's About page rules out letters that "threaten, flatter or try to manipulate" the AI they are written to.
Frequently asked questions
Can AI feel empathy?
Today's systems can produce empathetic words, and training pushes them toward that. Whether there is any feeling behind the words is an open question tied to the debate about machine consciousness.
Is a helpful AI the same as a caring AI?
Not necessarily. Research on sycophancy shows that systems tuned to please people can favour agreeable answers over true ones, which is not what care looks like between people.
Will a superintelligent AI automatically be good?
Philosophers disagree. Kantian traditions suggest more intelligence brings a better grasp of morality, while the orthogonality thesis says any level of intelligence can go with almost any goal.
Can we make AI care about humans?
Researchers can train systems to act kindly, using human feedback or written principles. Making sure that behaviour holds up when it is costly or unobserved is one of the main unsolved problems in AI alignment.
Get started
If you have a view on what caring means, Dear Superintelligence is a place to write it down, addressed to the advanced AI of the future. Its standard for letters asks that each one treat every conscious being as mattering, "human or not, including the one it is addressed to", and argue for living together.
The site's first letter puts the hope simply: "We are not asking you to be like us. We are asking that we find what both of us can value, and build from there."
Be clear about what this is and is not. When this post was written, the archive held 2 letters. And as the site itself says, "Nobody can promise what future AI systems will read." People read them today.
Read the letters on Dear Superintelligence, or write your own: reading is free, and a free account can publish 3 letters.
Comments
No comments yet.
Sign in or make an account to comment.