ai-labor-market

AI Voice Interviews Beat Human Recruiters in a 70,000-Applicant Trial

Only 15% of recruiters expected it. Randomly assigned AI voice interviews raised job offers 12% and starts 18%, no productivity loss. What it does not show.

ParÉditeur et auteur
Publié:
Analyse assistée par IA

Recruiters at a Philippine call-center staffing firm were asked to predict what would happen if an AI voice agent ran their job interviews instead of them. Only 15 percent thought applicants would get more offers. Then 70,884 applicants were randomly assigned to a human or to the AI, and the applicants interviewed by the machine were 12 percent more likely to be offered the job, 18 percent more likely to start it, and about as likely to still be there four months later. The recruiters were wrong about the direction. This post is about why, and about what the result does not say.

What was tested, and by whom

[Fact] The paper, "Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews," is a working paper by Brian Jabarian (Carnegie Mellon University) and Luca Henkel (Erasmus University Rotterdam). The version I read is v2, posted to arXiv on 11 September 2026 and dated 10 September. It is a preprint, not a peer-reviewed publication. The experiment was pre-registered on the AEA RCT Registry and approved by the University of Chicago IRB.

[Fact] The setting is PSG Global Solutions, a recruitment process outsourcing firm and subsidiary of Teleperformance, hiring entry-level customer service representatives in the Philippines for large US and European clients. The jobs pay between 16,000 and 25,000 Philippine pesos a month, which the authors convert to roughly 280 to 435 US dollars. Applicants who passed an initial screen were randomly assigned to be interviewed by a human recruiter, by an AI voice agent, or, in a third arm, given the choice. Interviews ran 10 to 20 minutes, in person or by phone, following the same written guidelines in both arms: up to 14 topics, a recommended order, sample questions, and freedom to adapt.

The design detail that matters most: [Fact] in every arm, the hiring decision was made by a human recruiter who reviewed the interview and a standardised language-and-reasoning test. The AI collected information. It did not decide. That separation is the whole point of the paper, and it is also the reason the headline should not be read as "AI replaces recruiters."

There is a disclosure you should know about. [Fact] Data collection finished on 7 June 2025. On 1 August 2025, Jabarian accepted an unpaid role as Chief Economist at PSG Global Solutions, renewing the research partnership for five years. The authors state the firm supplied the data but had no role in the analysis or the decision to publish. I take the disclosure at face value; I also note that the firm now has a standing relationship with the lead author and a commercial interest in the finding.

The numbers, as the paper reports them

[Fact] Applicants interviewed by a human received a job offer in 8.70 percent of cases. Applicants interviewed by the AI voice agent received one in 9.73 percent of cases. The paper reports the difference as 0.91 percentage points with p below 0.001, and as a 12 percent higher likelihood. [Estimate] Working from the two rates directly, 9.73 divided by 8.70 gives 1.118, so "12 percent" is a rounded relative figure and the absolute gap is about one applicant in a hundred. Both framings are true; only the relative one makes a headline.

[Fact] Downstream of the offer, applicants in the AI arm were 18 percent more likely to start the job, and 19 percent, 18 percent and 19 percent more likely to still be employed after two, three and four months respectively. [Estimate] Because the job-start gain (18 percent) is larger than the offer gain (12 percent), a rough division suggests that, conditional on getting an offer, AI-interviewed applicants were around 5 percent more likely to actually show up. That is my arithmetic on two published ratios, not a figure the paper reports, and it ignores differences in the denominators.

[Fact] On productivity of those hired, the authors examine three firm metrics, including customer volume handled and employer quality-assurance scores, and find no difference between hiring modes. [Fact] Applicants rated stress, comfort, follow-up fluency and feedback quality similarly across arms, with one exception: AI interviews were rated significantly less natural.

Why the machine did better: "controlled variance"

The mechanism the authors propose is not that the AI is smarter. It is that it is more consistent without being rigid.

[Fact] Using the transcripts, they measure how closely each interview followed the written guidelines. The average similarity between the questions actually asked and the guideline questions was 0.59 in AI interviews and 0.43 in human interviews. [Estimate] That is about 37 percent higher adherence. The AI also showed significantly lower variance in topic order. Neither the humans nor the AI read from a script; the difference is that human recruiters drifted more, and drifted differently from one another.

[Claim] The authors' argument is that when a firm delegates the same information-gathering task to many people, repeated thousands of times, the variation in how each person does it becomes noise in the signal the firm is trying to read. Structure removes the noise but also removes follow-up questions. An AI agent, they argue, can hold to the structure and still adapt to the individual, and the transcripts show AI-interviewed applicants displaying more of the communication signals that predict offers in human interviews too. This is a plausible reading of the evidence. It is also the paper's interpretation, and the transcripts are correlational once you move past the random assignment.

The applicants preferred the AI, which is not the same as the AI being better

[Fact] In the arm where applicants could choose, 78 percent chose the AI voice agent. Survey evidence in the sample shows generally positive attitudes toward AI, and those attitudes predicted the choice. [Fact] But the authors also find negative sorting: applicants who chose the AI had significantly lower measured quality than those who chose the human. So the 78 percent tells you what people wanted, not that self-selection into AI interviews improves a firm's hiring pool. A firm that offers the choice may get a different applicant mix than a firm that assigns.

What this does not show

First, the population. These are entry-level, high-volume, English-language customer service jobs in a country where the industry employs more than 1.5 million people, with structured 10-to-20-minute interviews. Nothing here speaks to a senior hire, a technical interview, or a role where the interviewer's judgment is the product. The "controlled variance" argument is explicitly about scale and repetition; it weakens as those disappear.

Second, the outcome window. Four months of retention and three productivity metrics is the horizon. Whether AI-interviewed hires diverge later is not in the data.

Third, "offer rate went up" is a firm-side outcome. The paper frames it as better information collection, and the retention and productivity results support that. But a 12 percent higher offer rate could also be read as the human evaluators trusting a cleaner transcript more, which is an evaluator effect rather than an applicant-quality effect. The authors control for test scores and find the interview-score effect survives, which cuts against this, but does not settle it.

Finally, the counter-narrative worth stating plainly. The tempting takeaway is that AI is coming for recruiters. In this experiment, every offer was made by a human recruiter, the same human recruiters, reviewing an interview someone else conducted. What the AI absorbed was the twenty-minute screening conversation, which is the part of a recruiter's job most likely to be handed off at a firm running tens of thousands of interviews a year. If you work in that role, the relevant page here is human resources specialists; if you are the person being interviewed for these jobs, it is customer service representatives. The study says the applicant did not lose from the switch. It does not say what the screening recruiter's job looks like afterwards.

Sources

  • Jabarian, B. and Henkel, L. (2026), "Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews," arXiv:2607.28222v2 [econ.GN], 11 September 2026 (paper dated 10 September 2026). https://arxiv.org/abs/2607.28222
  • AEA RCT Registry #15385; University of Chicago IRB #IRB24-1894 and #IRB25-1002, as stated in the paper.

Quotations are brief and used for commentary. The paper is a preprint and has not been peer reviewed.

AI-assisted analysis: this post was drafted with AI assistance and reviewed by a human editor before publication. Figures marked [Estimate] are this site's own calculations from the paper's reported numbers, not results the authors report.

Analysis based on the Anthropic Economic Index, U.S. Bureau of Labor Statistics, and O*NET occupational data. Learn about our methodology

Historique des mises à jour

  • Publié pour la première fois le 15 septembre 2026.
  • Aucune mise à jour substantielle depuis la publication initiale.

Tags

#arxiv#field-experiment#ai-interviews#hiring#recruiters#customer-service#philippines

Sources

  1. arxiv.org