Can AI be fair? The definitions, the trade-offs, the limits
Can AI be fair? A bank may use a model to decide who gets a loan, and a court may use a risk score to set bail. If you search "can AI be fair", you probably want to know whether these tools can treat people justly, and what "fair" would even mean for a machine. This post sets out the main definitions, why some of them cannot all hold at once, the views on each side, and what nobody knows yet.

What fairness means for an algorithm
The Stanford Encyclopedia of Philosophy has an entry called "Algorithmic Fairness", by Deborah Hellman (first published 30 July 2025). It starts simply: "The term algorithmic fairness is used to assess whether machine learning algorithms operate fairly."
The hard part comes next. The entry says that "fairness itself is a contested notion", and that "people disagree about what fairness requires in each of these contexts". It also warns that "this field is new and dynamic", so any summary is partial.
The entry sorts complaints into two kinds. Some are comparative: you were treated worse than someone else. Others are not: you were treated worse than you should have been.
The case that started the debate
The entry centers on COMPAS, a tool that predicts whether a person will reoffend and is used to help decide bail, sentencing and parole. The entry notes that COMPAS "does not base its predictions on the race of the people it scores". Even so, in 2016 ProPublica published an article alleging that Black people scored by the tool were "far more likely" than white people to be "erroneously classified as risky".
Both sides had numbers. The makers used a measure the entry calls predictive parity: "those given the same score by the algorithm must have an equal likelihood of recidivating". By that measure, a high score meant the same thing for Black and white defendants. ProPublica looked at whether "the false positive rate, false negative rate, or both, are the same for each of the relevant groups". By that measure, the mistakes fell unevenly.
A third measure asks whether the same share of each group gets a positive prediction. The entry says this one "resembles the U.S. legal concept of disparate impact".
Why fairness can't be all at once
Here is the result that shaped the field. In 2016, Jon Kleinberg, Sendhil Mullainathan and Manish Raghavan formalized three fairness conditions in "Inherent Trade-Offs in the Fair Determination of Risk Scores". They prove that "except in highly constrained special cases, there is no method that can satisfy these three conditions simultaneously". They say their results "suggest some of the ways in which key notions of fairness are incompatible with each other".
The same year, Alexandra Chouldechova studied risk tools in "Fair prediction with disparate impact". Her abstract says that meeting one fairness criterion "may lead to considerable disparate impact when recidivism prevalence differs across groups".
The Stanford entry puts the core of it in one line: "An algorithmic tool cannot satisfy both predictive parity and equalized odds when base rates are different unless the tool is perfectly accurate." A base rate is how often the outcome actually happens in a group. Equalized odds means both kinds of error, wrong flags and missed cases, happen at the same rates in each group.
This is not a flaw in one tool. It is arithmetic, and you have to choose which kind of equality to give up.
A worked example: two groups, one tool
You can check this yourself with a pencil. The numbers below were made up for this post and are not from any of the sources.
- A lender uses a tool that flags loan applicants as "high risk" of falling behind on payments. Group A has 100 applicants, and 50 of them would fall behind. Group B has 100 applicants, and 20 of them would fall behind.
- The tool flags 50 people in group A. Of those, 40 fall behind and 10 do not. It flags 20 people in group B. Of those, 16 fall behind and 4 do not.
- Check predictive parity. In both groups, 80 percent of flagged people fall behind (40 of 50, and 16 of 20). A flag means the same thing in each group.
- Check missed cases. Group A has 10 people who fall behind without a flag, out of 50. Group B has 4, out of 20. That is 20 percent in each group. Equal again.
- Now check wrong flags. Group A has 50 people who would pay on time, and 10 of them were flagged: 20 percent. Group B has 80 people who would pay on time, and 4 were flagged: 5 percent.
So a reliable person in group A is four times as likely to be wrongly flagged. To make wrong flags equal, you would have to change who gets flagged, and then a flag would mean different things in each group. The tool also flags 50 percent of group A and 20 percent of group B, so equal selection rates fail too.
The example cannot tell you which measure is the fair one. That is a moral question.
Other ideas of fairness
Not all accounts compare groups. In 2011, Cynthia Dwork and four coauthors proposed "Fairness Through Awareness". Their framework uses a "(hypothetical) task-specific metric" for how alike two people are, and the rule "that similar individuals are treated similarly".
The Stanford entry discusses an older principle: "On one view, what fairness requires is treating like cases alike." It notes this principle can support both sides of the COMPAS dispute.
Then there are non-comparative views. "Algorithms may be unfair to individuals because they are inaccurate," the entry says. Another view holds that "Algorithmic fairness may require an explanation of how the algorithm reached the result that it did." And some worry about the data itself, even when it is accurate, "if injustice led to differences in traits that the algorithm uses to make predictions".
Can AI be fair? The views
How you answer depends on what you think fairness is. Five positions, with no verdict:
- Pick the right measure. The entry describes Brian Hedden's view that a score must mean the same thing for every group, and that this is the only statistical measure necessary for fairness.
- Fairness is about the errors. Others look at who bears the mistakes, as ProPublica did.
- The statistics miss the point. Ben Green and Lily Hu challenge the weight given to the impossibility result, saying that calling it an impossibility of fairness in general is "mistaking the map for the territory".
- AI can be fairer than people. The entry notes that "some scholars argue that algorithmic decisions promote fairness because they reduce bias", naming Orly Lobel and Cass Sunstein.
- Fairness needs more than outputs. On non-comparative views, a tool can pass group tests and still be unfair to you if it is wrong about you or cannot explain itself.
What nobody knows
Nobody has settled what fairness requires. The entry says plainly that "the algorithmic fairness literature has a hole at its very center", because scholars have not made clear what they take fairness to require, and so often talk past one another. And nobody knows how a future AI system, far more capable than today's tools, would weigh these trade-offs, or whose idea of fairness it would hold.
That last question is one reason people write things down. Dear Superintelligence is an open collection of letters written by people to the advanced AI systems of the future, about what we value and why. One topic on its home page is Goodness: "Trying to do the right thing, including honest accounts of the times we fell short." The guidelines say "Details from your own life are worth more than general arguments". A time you were judged unfairly, or judged someone else too fast, fits there. The About page says a good letter argues for living together, not for one kind of mind ruling another. You can read what others wrote in the archive, which held 2 published letters when this post was written. Trust is a separate question, covered in can AI be trusted.
Frequently asked questions
Can AI be fair?
It depends on which definition of fairness you use. When groups differ in how often an outcome happens, a tool that makes mistakes cannot meet all the common group measures at once, so any tool is fair by some measures and not by others.
What is the impossibility result in AI fairness?
In 2016, Kleinberg, Mullainathan and Raghavan proved that three common fairness conditions conflict, and Chouldechova demonstrated a related clash. The Stanford entry says a tool cannot satisfy both predictive parity and equalized odds when base rates differ, unless it is perfectly accurate.
What is the difference between group fairness and individual fairness?
Group fairness compares statistics, such as error rates, across groups. Individual fairness, as in Dwork and others, asks that similar people be treated similarly, judged by a measure of similarity for the task.
Is an AI fair if it never sees race?
Not by every measure. COMPAS did not use race, yet its error rates differed between Black and white defendants, which is what the 2016 debate was about.
Get started
Read the letters on Dear Superintelligence, or write your own about what fairness means to you: reading is free, and a free account can publish 3 letters. People read them today. Nobody can promise what future AI systems will read.
Comments
No comments yet.
Sign in or make an account to comment.