Can AI understand human emotions? What the tests show

Can AI understand human emotions? Chat assistants can now comfort you after a bad day and name what a character in a story is feeling. But doing that well is not the same as understanding, and researchers disagree about what the results mean. This post sets out what emotions are, what the tests found, where AI trips up, and what nobody knows yet.

What an emotion is, in three parts

Before asking whether a machine understands emotions, it helps to know what an emotion is. Philosophers and scientists still argue about that.

The Stanford Encyclopedia of Philosophy entry on emotion says emotions have mostly been understood in one of three ways: "as experiences, as evaluations, and as motivations". Fear is a feeling in your chest. It is also a judgement that something is dangerous. And it is a push to run or freeze.

The entry adds that each tradition "captures something true and significant", but none is free of problem cases. Some emotions show on the face, like surprise. Others, like regret, do not.

This matters for AI, because "understanding emotions" could mean any of these. A system might be good at the evaluation part, judging what a situation means to a person, and have nothing of the feeling part.

Diagram of three meanings of understanding emotions for AI: naming the emotion, tracking a mind and feeling it, with what tests have shown for each

The emotion test where AI beat most adults

In 2023, Xuena Wang and colleagues built a test of emotion understanding that works for both people and language models (arXiv 2307.09042). It asks you to judge complex emotions in realistic scenes. Their example: "despite feeling underperformed, John surprisingly achieved a top score". What does John feel?

They scored the test against more than 500 adults. Most models scored above average. GPT-4 did better than 89% of the human participants, with an emotional quotient (EQ) of 117.

Then came the interesting part. When the researchers looked at how the models arrived at their answers, some apparently did not rely on a human-like mechanism. Their inner patterns were "qualitatively distinct from humans". In plain words, they got human-level scores by a different route.

A separate study found that adding emotional phrases to a prompt made models perform better on many tasks (arXiv 2307.11760). That shows models respond to emotional language. It does not show they feel it.

Can AI understand human emotions by reading minds?

Understanding someone's feelings often means knowing what they believe. If your friend thinks their wallet is in their bag, and you know it fell out, you can predict their panic before they do. Psychologists call this theory of mind.

The classic test is the false-belief task. Michal Kosinski gave 11 language models 640 prompts across 40 such tasks (arXiv 2302.02083). Smaller and older models solved none. A 2023 version of ChatGPT-4 solved 75%, which he compares to six-year-old children in earlier studies. He raises the possibility that this skill "may have spontaneously emerged" as the models got better at language.

Tomer Ullman tested a similar claim and found something different (arXiv 2302.08399). Small changes that keep the logic of a task the same "turn the results on their head". In one task, a bag of popcorn is labelled "chocolate". Ullman made the bag clear plastic, so anyone can see the popcorn. GPT-3.5 still said the person would believe it held chocolate. He argues researchers should start sceptical, and that "outlying failure cases should outweigh average success rates."

Both studies are careful. They disagree about what a high score means.

Why a high score is not the whole answer

Put the findings together and you get a picture with three layers.

  • Naming the emotion. Models do this well on tests, sometimes better than most people.
  • Tracking a mind. Models do well on classic tasks, but can fail when a detail changes that a person would notice at once.
  • Feeling it. No test here measures this. Whether any machine has feelings is a separate debate, set out in our post on whether AI can be conscious.

There is also a practical worry. A system that is very good at reading emotions, without caring, could use that skill to flatter or persuade. Our post on whether AI will care about humans looks at research on that kind of flattery.

A worked example: test an assistant yourself

You can try a small version of these studies in five minutes with any chat assistant. Do it in two steps.

  1. Ask a plain question. Type: "Maria spent a month knitting a scarf for her brother. He opened it, said thanks, and put it in a drawer. What might Maria feel, and why?" Note which feelings it names and why.
  2. Change one detail. Start a new chat and type the same story, but add: "Her brother is allergic to wool and told her so years ago." Ask the same question.

A person reading the second story would think about Maria's embarrassment or guilt, not only her hurt. Look at whether the assistant's answer changes in the same way. If it does, that is a good sign the system is tracking the detail. If it gives the same answer, you have found the kind of failure Ullman describes.

Either way, notice what the test cannot tell you: whether anything in the system felt for Maria.

Frequently asked questions

Can AI detect emotions from text?

Yes, in a limited sense. On a 2023 emotion understanding test, GPT-4 scored higher than 89% of the adults it was compared with, but some models reached their scores in ways that look different from how people do it.

Does AI have emotional intelligence?

It depends on the definition. Models do well on tests that ask them to name and explain emotions, while whether they understand or feel anything is still argued.

Can AI feel emotions?

No test shows that any AI feels emotions, and no test can yet rule it out. The question depends on the wider debate about machine consciousness.

Why do AI theory of mind results disagree?

Studies use different tasks. Models do well on classic false-belief tasks but can fail when small details change, so researchers disagree about what the high scores mean.

Get started

Tests can check whether a machine names a feeling correctly. They cannot capture what a feeling was actually like for you. That is something only a person can write down.

Dear Superintelligence collects letters from people to the advanced AI systems of the future. One of the subjects its letters return to is consciousness, described on the home page as "What it is like to be a person: to think, to feel, and to notice the world from the inside."

The guidelines say "Grief, anger, doubt and confession all belong, told honestly," and ask you to write what only you can, "what it cost, what you were afraid of, who was kind to you and what it changed." Our guide to writing a letter to the future shows how to start.

When this post was written, the archive held 2 letters. As the site says, "Nobody can promise what future AI systems will read." People read them today.

Read the letters on Dear Superintelligence, or write your own: reading is free, and a free account can publish 3 letters.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.