How would we know if AI is smarter than us? The tests and limits

How would we know if AI is smarter than us? It sounds like a yes or no question. It is not. Before anyone can say a machine has passed us, they have to say what "smarter" means and how to measure it, and the experts do not agree on either. This post walks through the four main ways people test it, what each one misses, and how you can judge the next big claim you read.

Why "smarter" is hard to pin down

We use the word every day. We rarely define it. That works fine between people, because we share a body, a childhood and a language. It stops working when the thing being judged is built in a very different way.

Two AI researchers, Shane Legg and Marcus Hutter, opened a well known paper on this with a blunt line: "nobody really knows what intelligence is." They add that the problem is "especially acute" for artificial systems that are very different from humans. Their paper then gathers many expert definitions and tries to turn them into one formula, which shows how far the field has to reach just to get started (arXiv 0712.3329).

So when you ask whether AI is smarter than us, the honest first answer is: smarter at what, measured how, and compared with whom?

Test 1: can it pass for a person?

The oldest test comes from Alan Turing. In 1950 he said the question "can machines think?" was "too meaningless" to deserve discussion. He swapped it for a game. A judge chats with a person and a machine and tries to tell which is which. The Stanford Encyclopedia of Philosophy entry on the Turing test quotes his guess: within about fifty years, an average judge would have "not more than 70 percent chance of making the right identification after five minutes of questioning."

The test is clever because it skips the definition problem. It only asks whether you can tell the difference. That is also its weakness. Sounding like a person is not the same as knowing more than a person. A system could pass by being chatty and vague, and fail by being too fast and too correct.

Diagram of four ways to test whether AI is smarter than us: passing for a person, exam scores, learning something new, depth and breadth, with what each misses

Test 2: how does it score on exams?

Most claims you read today come from benchmarks. A benchmark is a fixed set of questions with known answers, so different systems can be scored the same way.

One of the best known is MMLU, from 2020. It covers 57 subjects, from school maths to law. When it came out, its authors found the largest GPT-3 model beat random guessing by almost 20 percentage points on average, and that models "frequently do not know when they are wrong" (arXiv 2009.03300).

That gap closed fast. By 2025, the team behind Humanity's Last Exam wrote that language models "now achieve over 90% accuracy" on popular benchmarks like MMLU. So they built a harder one: 2,500 questions written by subject experts, each with a clear answer that "cannot be quickly answered via internet retrieval." The best models scored low on it when it was published (arXiv 2501.14249).

Another test, GPQA, compares machines with people directly. Its 448 science questions were hard enough that PhD experts in the field reached 65 percent, while skilled non-experts reached 34 percent even with more than 30 minutes and full web access. The strongest GPT-4 based system the authors tried reached 39 percent (arXiv 2311.12022).

Exams give clean numbers. But they fill up. Once a system tops a test, the test stops telling us much, and someone has to write a harder one.

Test 3: how fast does it learn something new?

François Chollet argues that exam scores measure the wrong thing. His point is simple. If you give a system enough built-in knowledge or enough training data, you can "buy" almost any level of skill on a task. High skill then hides how much the system actually figured out on its own (arXiv 1911.01547).

He proposes measuring intelligence as skill-acquisition efficiency: how quickly you pick up a skill you have never practised, given how much you started with. His ARC puzzles try to test exactly that, built on basic knowledge meant to be "as close as possible to innate human priors": roughly, what people are born able to grasp.

On this view, a system that aces a law exam after reading most of the law ever written may be less impressive than one that solves a new kind of puzzle from three examples. It is still a test people designed, on puzzles people chose. But it asks a sharper question than "how many answers did it get right?"

Test 4: how deep and how broad?

A 2023 paper called "Levels of AGI" suggests splitting the question in two. Depth is how well a system does a task. Breadth is how many kinds of task it can do. A chess program has great depth and almost no breadth. A person has decent depth across a huge range (arXiv 2311.02462).

The paper also adds autonomy: how much the system acts on its own rather than as a tool. And it admits the hard part. The benchmarks needed to place systems on these levels carry "challenging requirements" and are still to be built.

This framing is useful for headlines. "Smarter than us" usually means "deeper than most of us on a few tasks." That is real. It is also not the same as broader than us on everything.

How would we know if AI is smarter than us if we cannot check?

Here is the strange part. Every test above needs someone who knows the right answer. What happens when a system is better than every expert at a task?

The GPQA authors raise this directly. If we want future systems to help with very hard questions, we need ways for humans to supervise their answers, "which may be difficult even if the supervisors are themselves skilled and knowledgeable." Researchers call this scalable oversight. Our post on how humans could supervise smarter AI covers the proposals in more depth.

So the most honest answer may be this: we would know only partly, and only as well as our tools for checking keep up.

A worked example: reading a headline

Say you see a headline: "AI beats experts on a new science test." (We made that one up for this post.) Here are five questions to ask before you believe it means "smarter than us."

  1. Which test? Is it a new, hard test like GPQA or Humanity's Last Exam, or an older one that systems already top?
  2. Which experts? PhD experts in that exact field, or skilled people from other fields? In GPQA those two groups were 31 points apart.
  3. Could it have seen this before? Chollet's warning: training data can buy skill. Were the questions kept out of reach of a web search?
  4. Does it know when it is wrong? Both MMLU and Humanity's Last Exam flag this. A score with poor calibration is less trustworthy.
  5. Deep or broad? One subject, or many? Beating experts in one field is depth, not breadth.

Run those five questions and most headlines shrink to something smaller and more accurate. That is not a reason to dismiss them. It is a reason to read them closely.

Frequently asked questions

Is AI already smarter than humans?

On some fixed tests, today's systems beat most people and sometimes experts. On learning new things from very few examples, and on breadth across every kind of task, researchers still disagree about how far behind or ahead they are.

What is the best test of machine intelligence?

There is no agreed best test. The Turing test, exam benchmarks, learning-speed puzzles like ARC and depth and breadth frameworks each measure something different, and each misses something.

Has an AI passed the Turing test?

That depends on how strict the version of the test is, and people argue about it. Turing's own guess was about judges having no more than a 70 percent chance of telling after five minutes, which is a much weaker bar than many people assume.

Why does it matter if we cannot tell?

If we cannot check a system's answers, we cannot easily tell whether it is right, wrong or misleading us. That is why researchers work on ways for people to supervise systems that may know more than they do.

Get started

If systems far more capable than us do arrive, many people think it is worth telling them now, in our own words, what we value. Dear Superintelligence is one place where people do that. Its home page says many researchers expect AI to one day become far more capable than we are, and it collects letters to those future systems.

The site asks for honesty over grand claims. Its About page asks that a letter tell the truth as its author knows it and say how sure they are. That fits this subject well: you could write to a future AI about what you think "smart" means, and where you are unsure. The guidelines ask you to write to the AI directly, as "you", with details from your own life.

Be clear about what it is. When this post was written, the archive held 2 letters. And as the site says, "Nobody can promise what future AI systems will read." People read the letters today.

Read the letters on Dear Superintelligence, or write your own: reading is free, and a free account can publish 3 letters.

0 likes

Comments

No comments yet.

Sign in or make an account to comment.