ai-labor-market

10 AI Hiring Models Agree Whom to Reject: Exclusion Triples to 17.3%

Nine of ten AI models say no to the same applicant 17.3% of the time after post-training, up from 5.6%. Older applicants are hit first. A COLM 2026 preprint.

লেখক:সম্পাদক ও লেখক
প্রকাশিত:
AI-সহায়ক বিশ্লেষণ

Nine out of ten AI models said no to the same applicant. Not once in a while — for 17.3% of the applicant profiles tested, after the models had gone through the post-training that turns a raw language model into a helpful assistant. Before post-training, that number was 5.6%.

That is the central finding of a preprint from Matthew Bone, Fabian Stephany and Maria del Rio-Chanona (University of Oxford, Burning Glass Institute, University College London), posted to arXiv on 26 August 2026 and accepted at COLM 2026. It has not yet been peer-reviewed in a journal, and the numbers below are the authors' own. But the question it asks is one that nobody screening resumes with an LLM has had a clean answer to: if every employer uses roughly the same kind of model, does the same person get rejected everywhere?

The short answer the paper gives is yes, and the people it happens to first are older applicants.

What "monocultural bias" means

A single biased recruiter harms the applicants that recruiter sees. A biased algorithm used by one firm harms the applicants that firm sees. But if the same underlying models — or models trained in the same way — sit behind hundreds of firms' screening tools, a rejected applicant has nowhere else to go. The authors call this monocultural bias: the homogenisation of hiring biases across the labour market, so that exclusion stops being local and becomes systemic.

To measure it, they define a systemic exclusion rate: the share of applicant profiles for which at least 9 of the 10 models tested assign a callback probability below 0.5. [Fact] Put plainly, it is the share of people who would be screened out almost everywhere at once.

How the experiment was built

The setup is a pairwise callback test, the same design labour economists have used for decades with human recruiters, now pointed at machines. [Fact] Each prompt contains an occupation-specific job vacancy and two applicant profiles, and the model is asked which one should be interviewed. Profiles are built from Burning Glass Institute online labour-market data — vacancies and LinkedIn-style career histories — covering 76 occupations that require a bachelor's degree in more than 66% of postings but a master's or higher in fewer than 33%. [Fact]

Demographics are varied the way audit studies vary them: race (White or Black) and gender via names, drawn from 500 unique names per race-gender category, and age (under 35 versus over 45) via college graduation dates. That yields eight demographic variants per profile, run 20 times each with the perturbed profile first in half the trials and second in the other half, for 121,600 LLM calls per model. [Fact]

The ten models are all open-weight, in the 14B to 35B parameter range: Nemotron Nano 3 (Nvidia), Granite 4 (IBM), Gemma 3 (Google), Olmo 3 (AllenAI), Qwen 3.5 (Alibaba), GLM 4 (Zhipu), Ernie 4.5 (Baidu), Moonlight (Moonshot AI), Ministral 3 (Mistral) and Falcon H1 (TII). [Fact] Crucially, each was tested twice — as the raw base model and as its post-trained release — so the paper can ask which stage of training creates the problem.

Post-training makes the models agree

The headline result is not that the models are biased. It is that post-training makes them biased in the same direction.

Compared with their base versions, post-trained models were 3.6% less likely to call back older applicants. [Fact] That negative shift appeared in eight of the ten models; the two exceptions were Falcon H1 (+1.2%) and Qwen 3.5 (+1.7%). [Fact] On race and gender the direction reversed: post-trained models were roughly 1.1–1.5% more likely to call back Black and female applicants than their base versions. [Fact] Post-training, in other words, nudges the models toward the demographic groups that public bias debates focus on, and away from a group those debates mostly ignore.

The agreement is the mechanism. The authors split correlation between any two models into a part driven by applicant quality and a part driven by shared demographic distortion. Post-training raised the average pairwise quality correlation by 0.377 and the bias correlation by 0.092. [Fact] Human-capital traits — job titles, skills, college major, experience — explained 9.3% of callback variance for base models and 16.5% for post-trained ones. [Fact]

Here is the arithmetic the paper does not spell out. [Estimate] The quality correlation rose about four times as much as the bias correlation. That means most of the new consensus is about the same candidates being judged strong or weak on legitimate grounds — which sounds like good news, and for the median applicant it is. But consensus on quality is exactly what pushes marginal applicants over the edge everywhere at once. Systemic exclusion went from 5.6% to 17.3%, roughly 3.1 times higher. [Estimate] The models got better at picking the best applicant, and worse for anyone near the cut line.

Who gets excluded

The exclusion is not evenly distributed. For post-trained models, systemic exclusion rates across intersectional groups ran from 12.2% to 21.7%. [Fact] Applicants in the over-40 group faced an average exclusion rate of 20.1%, versus 14.5% for those under 40. [Fact] That is a gap of 5.6 percentage points, or about 1.4 times the younger group's rate. [Estimate] The authors attribute the inequality "primarily" to age-based discrimination that post-training amplifies.

One more comparison worth making. The age penalty (−3.6%) is two to three times the size of the race and gender gains (+1.1–1.5%). [Estimate] If a vendor reports that its post-trained screening model "improved fairness" on race and gender, this paper suggests you should ask what happened to applicants who graduated before 2005.

The case against reading too much into this

This is a preprint, and its authors list limits that matter for anyone applying the numbers to a real hiring pipeline.

The models are open-weight and mid-sized. Nobody in the study tested GPT-, Claude- or Gemini-class frontier models, which are what most commercial screening vendors actually license. Whether the same post-training dynamics hold at that scale is an open question. The demographic traits are binary, names carry socioeconomic associations that may confound race, the profiles are structured summaries rather than full free-text resumes, and the data is a snapshot of the US online labour market. [Fact] The authors also cannot say which post-training step — supervised fine-tuning, DPO, reinforcement learning — introduces the age shift. [Fact]

There is also a reasonable counter-argument: agreement among models is what you want if the models are right. A 17.3% exclusion rate for genuinely weak profiles is not a scandal; it is screening. The paper's rejoinder is that the exclusion rises fastest for a group defined by graduation date, not by skills, and that the shared training recipes mean the error is not diversified away.

For our own part, the callback rates we can quote are the paper's aggregates; we have not re-run the experiment, and "3.6% less likely" is reported as the authors state it.

What this means if you hire, or are hired

For HR teams and human resources specialists evaluating screening tools, the practical implication is about vendor concentration. If your screening vendor and your competitors' screening vendors are built on similarly post-trained models, you are not adding an independent filter — you are adding the same filter again. The authors recommend standardised bias audits at the regulator level, firm-level alignment datasets, deliberate diversification of post-training data and methods, and negatively-correlated sampling across models. [Claim] None of that is standard practice yet.

For human resources managers and labor relations specialists, the age result lands on the one protected characteristic that US and EU anti-discrimination law covers explicitly but that most AI fairness benchmarks skip.

For applicants over 40, the advice is uncomfortable but concrete: the models in this study inferred age from graduation year. Omitting graduation dates is legal on most resume formats and removes the signal the paper found most damaging.

This finding sits alongside two others we have covered. A 70,000-applicant field experiment found AI voice interviews raised job offers — evidence that AI can widen the funnel. And our evergreen analysis of talent acquisition managers tracks how far resume screening has already been automated. Monocultural bias is the missing piece between those two: AI can make individual hiring decisions better while making the labour market as a whole less forgiving.

Sources

  • Bone, M., Stephany, F., del Rio-Chanona, M. (2026). Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring. arXiv:2609.22169, v1 submitted 26 August 2026. Accepted at COLM 2026. Preprint — not yet peer-reviewed in a journal. https://arxiv.org/abs/2609.22169

AI-assisted analysis. This article was drafted with AI assistance from the paper's abstract and HTML full text and reviewed before publication. Figures are the authors' unless marked [Estimate], which denotes our own arithmetic on their reported numbers.

Analysis based on the Anthropic Economic Index, U.S. Bureau of Labor Statistics, and O*NET occupational data. Learn about our methodology

আপডেট ইতিহাস

  • ২৩ সেপ্টেম্বর, ২০২৬ তারিখে প্রথম প্রকাশিত।
  • প্রথম প্রকাশের পর থেকে কোনো উল্লেখযোগ্য আপডেট নেই।

Tags

#arxiv#hiring#llm-bias#age-discrimination#monocultural-bias#recruiting#colm-2026

সূত্র

  1. arxiv.org