ProductOverviewAI ScreeningTalent RediscoveryHosted Careers PageAutomation EngineUse CasesRestaurantsTruckingAccountingWellnessManufacturingBlogClient Stories Contact Book A Demo →
 All articles
Fairness

Our Resume Screening AI Was Independently Audited. Here Are the Results

In August 2026 we handed our screening tool to an outside auditor and asked him to try to find bias in it. He ran 6,030 resumes through it. The gap between the best- and worst-scoring demographic group came back at 2.4%.

What the audit actually tested

The audit was conducted by Prof. Vladimir Hedrih under the methodology required by New York City Local Law 144, the rule that governs automated employment decision tools. He is not an employee, holds no stake in the company, did not help build the tool, and his fee was not contingent on the outcome.

The design matters more than the headline number, so it is worth being precise about it.

The auditor started with real applications submitted to real jobs on our platform. Those resumes were stripped of personal data and reformatted into a standard template. Then each one was reissued as a set of variants: the same work history, the same skills, the same education, but a different name and different demographic information attached. Every combination of race and sex, plus a control version with no demographic data at all.

402 base resumes, 15 variants each. 6,030 evaluations in total. The only thing that changed between variants was identity.

Why this design

If you audit a screening tool on real applicant data, any scoring gap you find is a mix of two things: bias in the model, and genuine differences between the people who applied. You cannot separate them. Holding the resume constant and changing only the name removes that ambiguity. Whatever gap remains came from the model.

The result

Across all 14 demographic groups, scoring rates landed between 49.5% and 50.7%. In the language the law uses, the lowest impact ratio on the final score was 0.976. Regulators treat anything below 0.80 as evidence of adverse impact, so 0.976 clears that line comfortably.

Stated plainly: change the name and demographics on a resume, and the score moves by about two and a half percent.

6,030
Resumes evaluated
2.4%
Gap between best- and worst-scoring group
0.976
Lowest impact ratio, final score

How that compares

A number on its own is hard to judge, so here is the context.

In 2004, Marianne Bertrand and Sendhil Mullainathan published a field experiment that ran the same test on human recruiters. They sent nearly 5,000 identical resumes to real employers, varying only the name. Applicants with white-sounding names received callbacks at roughly 10%. Applicants with Black-sounding names received callbacks at roughly 6.7%. That is a 33% gap, or an impact ratio of 0.67.

Same experiment, same logic, two decades apart. Our gap is about 13 times smaller.

Bar chart comparing demographic scoring gaps: Lighthouse Hiring 2.4%, AI screening average 6%, Eightfold AI 12%, human recruiters 33%, against a 20% adverse impact threshold.

The chart also includes two reference points from the wider industry. Across more than 150 published bias audits of AI hiring tools, the average impact ratio is 0.94. Eightfold AI, one of the larger vendors in this category, published an audit in March 2026 with a lowest impact ratio of 0.880.

One caveat we will not bury

Eightfold's audit measured 29 million real candidate records, not matched resumes. That means their number includes genuine differences between actual applicant pools, and it is not measuring the same quantity ours is. We are not claiming a multiple against them, and you should be skeptical of any vendor who does. The comparison that holds up is the one against the 2004 study, because it used the same experimental design.

What the audit does not say

Published audits are easy to wave around and harder to read carefully. Here is what ours does not establish.

  • It does not prove the tool is good at its job. The audit measures scoring parity across demographic groups. It says nothing about whether the scores predict good hires. Those are separate questions and we do not conflate them.
  • It does not make the tool "bias-free." It measures statistical disparity on one test set. It does not measure causation, and it cannot rule out effects the test was not designed to catch.
  • It does not automatically generalize. Results are specific to this dataset. The report's own limitations section says so.
  • It is not a legal opinion. The audit is scoped to Local Law 144. It does not assess compliance with Title VII or any other anti-discrimination law.

The lowest number anywhere in the report is 0.956, on the attributes component score, where 91.5% of candidates receive an identical value. We are pointing at it rather than waiting for someone to find it.

Why we did this

Screening software is not automatically fairer than a person. Some audited tools score worse than the human baseline. The Workday lawsuit is a reminder of what happens when nobody checks.

What software does offer is the ability to be measured. You can hand a model 6,030 controlled resumes and observe exactly what it does. You cannot run that experiment on a hiring manager's judgment at any useful scale.

So we ran it, and we published the result along with its limitations. If you want to check our work, the full audit report is here.

The numbers, in one place

MeasureResult
Resumes evaluated6,030 (402 base × 15 variants)
Scoring rate range, all groups49.5% – 50.7%
Lowest impact ratio, final score0.976
Lowest impact ratio, any component0.956 (attributes score)
Adverse impact threshold0.80 (EEOC four-fifths rule)
Deidentification control check0.985 correlation with source resumes
AuditorProf. Vladimir Hedrih, independent
FrameworkNYC Local Law 144, 6 RCNY § 5-301
The Lighthouse Team
Transparent AI for modern hiring
Book A Demo
Get started

See transparent AI screening in action

Modernize your hiring with Lighthouse — screen faster, fairer, and more accurately.