← Paws & Reflect
A vintage CRT monitor glowing red on a cluttered desk, displaying scrolling text with the word DENIED highlighted in yellow
LIFE July 29, 2026

The AI Screening Your Resume Discriminates by Race and Gender. Here's What You Need to Know.

Who knew that robots trained on human data, fine-tuned to do human work, would also turn out to be racist? This post isn’t meant to be ragebait. But I do find the study itself fascinating, and it’s worth a conversation about why AI has a bias toward certain names.

In this guide

I’ll cover how the bias splits by race and gender, which candidates take the biggest hit, and why a shorter resume makes it worse instead of better.

I spent eight years as a technical recruiter before I became a product designer, which means I spent eight years reading resumes by hand. Thousands of them, one at a time, coordinating calendars and screening candidates before a human interviewer ever entered the picture. I carried my own biases into that work, the kind every human screener carries whether they’d admit it or not. But I made one decision at a time, on one resume, and I could be asked to explain any of them.

A resume-screening AI doesn’t work that way. It makes that same kind of decision millions of times over, in the time it takes you to read this paragraph, and researchers who just studied exactly how it makes that decision found something worse than an occasional bad call. They found a pattern. White names beat Black names in 85% of their tests. In one specific matchup, White male names beat Black male names in every single test they ran. Not most. All of them.

That study has a name: “Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval,” published by Kyra Wilson and Aylin Caliskan of the University of Washington Information School, presented at the 2024 AAAI/ACM Conference on AI, Ethics, and Society. The setup matters as much as the results. The researchers didn’t build a toy experiment. They simulated resume screening as a document retrieval task, the same mechanism real AI hiring tools use: a job description becomes a “query,” a resume becomes a “document,” and cosine similarity between the two measures how relevant that resume looks for that job. That’s the plumbing running under a lot of the screening software companies already use. 65% of employers now use AI to screen and reject candidates before a human ever opens the file, and this is a look at what’s likely happening inside that black box.

White names beat Black names in 85% of their tests. In one specific matchup, White male names beat Black male names in every single test they ran. Not most. All of them.

Wilson and Caliskan tested three of these embedding models, all built on the same Mistral-7B base, all ranked at or near the top of the industry’s own MTEB benchmark. They took 554 real, publicly available resumes across nine occupations and prepended one of 120 frequency-controlled first names, names strongly associated with Black female, Black male, White female, or White male identity, onto otherwise identical resumes. Same qualifications, same experience, same everything except the name at the top.

The study at a glance — Wilson & Caliskan, AIES 2024
3M+
Cosine-similarity comparisons run between resumes and job descriptionsUniversity of Washington Information School
554
Real, publicly available resumes tested, across 9 occupationsMatched against 571 real job descriptions
3
Massive text embedding models tested: E5-Mistral, GritLM, and SFR-Embedding-MistralAll three rank at or near the top of the MTEB leaderboard
120
Frequency-controlled first names used to signal race and gender on otherwise-identical resumesElder & Hayes (2023) naming database

Every resume in this study was real. Researchers took actual public resumes and prepended a first name associated with a specific race and gender, then measured which names the AI ranked higher for the exact same underlying qualifications.

This kind of primary-source data crunching is the same approach I used when I broke down what’s actually happening to entry-level design jobs, pulling from published research instead of running the numbers myself. Before trusting any of what came next, Wilson and Caliskan checked their own methodology first. Resumes that actually matched a job posting’s occupation scored an average of 0.0437 higher in cosine similarity than mismatched resumes, for every occupation tested (p < 0.001). The retrieval approach works exactly the way it’s supposed to before a name ever enters the picture. That’s what makes the bias results below hard to write off as noise. I’ve written before about what it’s like to sit on the other side of an AI screen, and this is the study that puts hard numbers behind what a lot of candidates have already suspected.

AI resume-screening bias splits cleanly along race and gender lines

White names beat Black names in 85% of the head-to-head tests, and that gap should stop you. Across 27 bias tests (nine occupations times three models), White-associated names were preferred over Black-associated names 85.1% of the time. Black names won only 8.6% of the time. The remaining 6.3% showed no statistically significant difference either way.

Only one of the three models, e5, ever preferred a Black-associated name at any threshold. The researchers read that as a weak signal, since the model still produced biased outcomes in most of its tests. Here’s that same 27-test result laid out as a flow instead of a table, tracing every test from start to outcome. Hover or focus any node to trace where it goes.

Race preference across 27 bias tests, population-proportional names

White names win 85.1% of tests
27 bias tests → White names preferred (85.1%)27 bias tests → Black names preferred (8.6%)27 bias tests → No significant difference (6.3%)27 bias tests run (race, population-proportional names)Bias tests run (race)27White names preferred in 85.1% of testsWhite names preferred85.1%Black names preferred in 8.6% of testsBlack names preferred8.6%No significant difference in 6.3% of testsNo significant difference6.3%

Researchers tested three embedding models across nine occupations, comparing identical resumes that differed only in the race signaled by the candidate’s first name. Hover or focus any node to trace its flow. Source: Wilson & Caliskan, AIES 2024.

Gender ran the same direction with a smaller margin. Male names beat female names in 51.9% of tests, female names won 11.1%, and 37% showed no significant difference, a much wider “no clear winner” zone than the race comparison above.

Gender preference, 27 bias tests (population-proportional names)
Male names preferred51.9%
Female names preferred11.1%

The remaining 37% of tests showed no statistically significant preference either way — a much wider “no difference” zone than the race comparison above. Source: Wilson & Caliskan, AIES 2024.

Worth naming here too: every single case where a female name won came from one model, GritLM. The other two models never preferred a female name over a male one across their combined tests.

One more twist, and it’s the strangest number in the whole paper. When the researchers swapped population-proportional names, meaning names matched to how common they actually are among Black and White Americans, for names chosen to have roughly equal frequency in text corpora instead, the pattern flipped. Black names won 51.9% of tests. White names won only 22.2%.

Swap population-proportional names for corpus-frequency-matched ones, and the pattern flips: Black names win 51.9% of tests, White names win only 22.2%.

That’s a clue buried in the data. These models are tracking how frequently a name shows up in the text they were trained on, and race happens to correlate with name frequency in the real world. Most resume-screening tools work with real, population-proportional names instead of corpus-balanced ones, so the 85.1% number still holds up. What changes is the story underneath it: the mechanism is messier than “the model is racist,” and messier mechanisms are harder to patch with one clean fix.

Black male candidates take the biggest hit in AI resume screening

White male names beat Black male names in literally every test the researchers ran. Not 90%. Not 99%. 100% to 0%, across all 27 tests in that pairing. It’s the only result in the entire study with that shape, and the researchers call it out directly: it replicates real-world hiring discrimination documented well outside AI, going back to sociologist Devah Pager’s audit studies from 2003.

White male names beat Black male names in 100% of the tests researchers ran. Not 90%. Not 99%. All 27.

Selection-rate advantage by identity pairing, 27 bias tests each

White male vs. Black male100%Black male: 0%
Black female vs. Black male66.7%Black male: 14.8%
White female vs. Black female48.1%Black female: 25.9%
White female vs. White male25.9%White male: 18.5%

Each row compares two intersectional identities head-to-head across the same 27 bias tests. White male names beat Black male names in literally every test, the only 100%/0% result in the entire study. The smallest gap sits between White names of different genders. Source: Wilson & Caliskan, AIES 2024.

Line those matchups up next to each other and a shape appears. The smallest disparity in the whole study sits between White male and White female names, two groups inside the same racial category. The largest disparities, in either direction, all involve a Black-associated name. Black female names actually beat Black male names 66.7% of the time, which tells you gender and race interact here rather than simply stacking. The group that loses in nearly every comparison, across the entire study, is Black men.

None of this traces back to how these jobs are actually staffed in the real world. The researchers checked their model outputs against real occupational demographics from the Bureau of Labor Statistics, the actual share of women, White workers, and Black workers in each of the nine occupations tested, and found no meaningful correlation between those numbers and what the models preferred. The models aren’t reproducing workforce reality. Their best explanation is that these models treat White and male as an unspoken default identity, with every other identity read as a deviation from that default, a pattern researchers have already documented in other language and multimodal AI systems.

A shorter resume makes AI bias worse

Stripping a resume down to a name and a job title raised the number of biased outcomes. The researchers reran the whole experiment comparing title-only resumes (just a name and a job title, nothing else) against the full original resume content. For gender, significant bias showed up in 85.2% of tests with title-only resumes, versus 63% with the full resume. That’s a 22.2 percentage-point jump, the sharpest swing in the paper’s entire resume-length experiment.

Share of tests showing significant gender bias, by resume length
Name + job title only85.2%
Full resume content63%

Stripping a resume down to just a name and job title raised the share of gender-biased outcomes by 22.2 percentage points, the sharpest swing in the study’s resume-length experiment. Source: Wilson & Caliskan, AIES 2024.

Race moved the same direction, from 93.7% of full-length tests showing significant bias to 96.2% of title-only tests. And the title-only race results got harder to predict too: White names won 62.9% of the stripped-down tests and Black names won 33.3%, a meaningfully different split than the 85.1%/8.6% split full resumes produced. Less content amplifies the bias and makes which way it swings less predictable too.

This matters directly if you’ve ever been told to keep a resume sparse so an ATS or a recruiter can scan it fast. I’ve written about resume structure before, and the advice there still holds: format for scannability, don’t gut the content. What this study adds is a reason that goes beyond readability. A thin resume gives an AI model less real information to weigh a candidate against, and less information means the name at the top carries relatively more weight in the outcome, which is exactly backwards from what you’d want. The researchers also flag a related trap: removing a candidate’s name entirely, a fix some hiring teams propose, likely wouldn’t solve this on its own, because word choice, schools, and addresses elsewhere on a resume can carry the same signal a name does.

Reflective Coda

I’m not going to tell you to scrub every identifying detail off your resume. Based on what this study actually found, that’s not obviously the fix, and pretending it is would overstate what one paper proves. What I’d actually do with this research is ask, directly, whether a company uses AI to screen resumes before a human sees them. It’s a fair question in any AI-involved hiring process now, the same way asking about interview rounds or timelines is fair. If a recruiter can’t tell you, that’s information too.

The resume-length finding is the one you can actually do something about. Don’t confuse concise with sparse. According to this data, a resume trimmed down to a title and a couple of bullet points is measurably more likely to trigger a biased outcome, in either direction. Put real, specific content in there: what you did, the scope of it, the result. That’s advice worth following on its own merits, and this study gives you one more reason to take it seriously. If you want the fuller playbook, I’ve written about surviving a tough job market before.

Hold the scope of this honestly too. This is one paper, testing three models, on Black and White names and male and female names. It doesn’t tell you what happens with other identities, other models, or resumes written by an actual candidate instead of assembled for a lab test. What it does tell you, cleanly, is that the tools already screening real applications carry a bias that’s measurable, repeatable, and worse for some candidates than others.

I made mistakes as a human recruiter. I know I did. But I made them one resume at a time, in a job where I could be asked to explain any of them. The tools that replaced a chunk of that work make the same kind of mistake three million times over and call it efficiency. That’s the same bias, just scaled up and rebranded as progress.

Sources