
"Is the AI biased?" is the wrong question, because it has no actionable answer. Bias is not a property a system has or lacks. It enters through specific mechanisms, and each one has a different test and a different fix.
There are three that matter.
Mechanism 1: the training data
Any system tuned on your past hiring decisions learns your past hiring decisions. Where those were sound, that is the point. Where they were skewed, the system now reproduces the skew consistently and at scale, with an audit trail that makes it look rigorous.
The canonical example is Amazon's experimental recruiting tool, abandoned in 2018. Trained on a decade of resumes submitted to a heavily male-dominated engineering organisation, it learned to downgrade resumes containing the word "women's" — as in "women's chess club captain" — and to penalise graduates of two all-women's colleges. Nobody encoded a preference for men. The system inferred it from what success had historically looked like.
The test: ask what the target variable is. If it is "resembles people we hired" or "resembles people who performed well according to past reviews", you have inherited whatever was in those decisions. If it is "meets the stated requirements for this role", you have not.
The fix: score against explicit requirements you wrote, not against a similarity-to-past-hires signal. This is the strongest argument for criteria-based screening over lookalike models, and it is a question worth putting to any vendor directly.
Mechanism 2: proxy variables
Remove name, gender and photo and you have removed the direct signals. You have not removed the correlated ones.
Postcode correlates with race and class in most countries with a history of residential segregation. University correlates with class and, in many places, with ethnicity. Employment gaps correlate with parenthood and disability. Continuous employment history correlates with health. Even writing style carries signal about whether English is a first language.
None of these are discriminatory to consider in principle — a university can be genuinely relevant. The problem arises when a proxy is doing work you have not noticed and would not defend if asked.
The test: for each input, ask whether you could justify it out loud to a rejected candidate. "We prioritised candidates with continuous employment" is a sentence most people would not want to finish.
The fix: drop inputs you cannot defend. For the ones you keep, check whether they are actually predictive or merely conventional — a lot of degree requirements survive on habit rather than evidence.
Mechanism 3: the ranking cutoff
The least discussed and, in day-to-day operation, often the most consequential.
Scoring systems produce a distribution. In a typical pipeline the top forty candidates might score between 71 and 78 — a spread comfortably inside the noise. But the output is an ordered list, and whoever reads it stops somewhere.
A difference too small to be meaningful becomes a hard boundary because it is rendered as a position. And if any group sits marginally lower on average for reasons unrelated to ability — because their resumes parse worse, because they use less assertive phrasing, because they took time out — a small average gap turns into a large difference in who gets read.
The test: compare selection rates at your actual cutoff, not average scores. Two groups can have nearly identical mean scores and very different pass rates if the cutoff falls in a dense part of the distribution.
The fix: work in bands. Everything above a threshold goes in one pile and the whole pile gets read. If the pile is too large, tighten a requirement — do not read further down the list.
Running an actual audit
The four-fifths rule is the usual starting point: if any group's selection rate is below 80% of the highest group's rate, investigate.
Concretely:
- Take real applicant data from a completed hiring round — not synthetic tests.
- Group by the characteristic you are checking, using whatever demographic data you legitimately hold. Where you hold none, geographic and name-based inference is unreliable enough that it is usually better to audit what you can and be explicit about the gap.
- Calculate the selection rate per group at your actual cutoff.
- Divide each by the highest rate. Anything under 0.8 is a flag.
- If flagged, work backwards through the three mechanisms above to find which one is producing it.
A flag is not proof of discrimination. It is a signal that something needs explaining — and if the explanation is a genuine job requirement, that is a defensible answer, provided you can show your working.
The compliance overlay
In several jurisdictions this has moved from good practice to obligation. NYC Local Law 144 requires an annual independent bias audit of automated employment decision tools, with results published. The EU AI Act classifies recruitment systems as high risk. Illinois, Colorado and Maryland have their own requirements.
The details, and what they mean in practice, are in AI hiring compliance in 2026.
The part worth keeping in view
The comparison that matters is not "automated screening versus a perfectly fair process." It is "automated screening versus a tired human reading application 250 at 6pm on a Friday."
Human screening is also biased, also inconsistent, and considerably harder to audit — you cannot run a four-fifths analysis on a recruiter's intuition. The advantage of an automated system is not that it is neutral. It is that it is measurable, which makes it fixable.
That advantage only pays off if someone actually measures.
Common inquiries regarding this topic.
How does bias get into AI hiring tools?
Through three distinct routes. Training data bias, where the system learns from past hiring decisions that were themselves skewed. Proxy variables, where a neutral-looking feature such as postcode, university or a career gap correlates with a protected characteristic. And ranking effects, where a scoring gap too small to be meaningful becomes a hard cutoff because results are presented as an ordered list.
Through three distinct routes. Training data bias, where the system learns from past hiring decisions that were themselves skewed. Proxy variables, where a neutral-looking feature such as postcode, university or a career gap correlates with a protected characteristic. And ranking effects, where a scoring gap too small to be meaningful becomes a hard cutoff because results are presented as an ordered list.
The Resume World Team
VerifiedProduct & hiring research, Resume World
We build the screening engine behind Resume World. Everything here comes out of working on resume parsing, scoring and hiring workflows day to day — including the parts that turned out harder than expected.
See more than just keywords.
Resume World extracts verifiable evidence from every applicant against role criteria and delivers an explained, ranked shortlist. 100% free to start with zero card required.



