AI Job Interviews: What 70,884 Applications Showed
Audience: Intermediate readers evaluating AI in recruitment, HR operations, or high-stakes workflows.
One of the largest field experiments on AI job interviews did not test autonomous hiring. It tested a narrower system: an AI voice agent conducted a structured interview, then a human recruiter reviewed the recording, transcript, and test results and made the decision.
That boundary produced a striking result. Among 67,056 eligible applications randomly assigned to an interview condition, the offer rate rose from 8.70% with human interviewers to 9.73% with AI interviewers—a 12% relative increase. Job starts and early retention also improved. Yet the AI workflow took longer from application to job start because human reviewers became the next queue.
The practical conclusion is specific: AI can improve repeatable information collection in high-volume hiring. The study does not show that an LLM should rank candidates, decide who advances, or replace accountable human review.
What did the AI job interview experiment test?
The Voice AI in Firms field experiment followed 70,884 applications received by a recruitment process outsourcing firm in the Philippines between March 7 and June 7, 2025. The jobs covered 48 postings, 41 client accounts, 19 cities, and sectors including technology, insurance, retail, finance, and healthcare.
The researchers randomized 67,056 eligible applications into three groups:
Both interviewers followed the same structured guidelines. The AI disclosed its identity at the start of the call. It also told applicants that a human recruiter—not the AI—would evaluate the interview and decide whether to make an offer.
That design matters more than the phrase “AI recruiter.” The AI collected evidence. Humans retained decision authority.
What did the researchers measure?
The paper measured outcomes beyond interview completion. It connected the randomized interview condition to offers, accepted offers, job starts, retention for up to four months, separation reasons, and available productivity measures.
| Outcome | Human interviewer | AI interviewer | Reported difference |
|---|---|---|---|
| --- | ---: | ---: | ---: |
| Job offer | 8.70% | 9.73% | 12% relative increase |
| Job start | 5.65% | 6.71% | About 18% relative increase |
| Employed after 30 days | 4.97% | 5.85% | About 17% relative increase |
| Median application-to-job-start time | 20 days | 24 days | AI workflow was four days slower |
The offer, start, and 30-day retention numbers use all randomized applicants in the human and AI interviewer groups. Later retention differences stayed positive, but estimates became less precise as the sample shrank. The authors also found no significant productivity difference in the subset of hired workers for whom productivity data was available.
These are causal estimates for this experiment because the researchers randomized the interviewer. They are not universal performance guarantees for every job, country, voice model, or hiring process.
Why did structured AI interviews collect better evidence?
The authors describe the mechanism as “controlled variance.” Human recruiters differed in question coverage, follow-up behavior, and the point at which they screened someone out. The AI agent followed the protocol more consistently while still asking responsive follow-ups.
Transcript analysis found that AI-led interviews covered more job-relevant topics and produced more information for the later human evaluation. This helps explain the higher offer rate without assuming that the model identified talent on its own.
Structured interviews already aim to ask comparable questions and score job-related evidence consistently. The AI agent made that operating discipline easier to repeat across thousands of calls. Teams considering AI job interviews should treat protocol adherence as the product requirement, not conversational realism alone.
The experiment also exposed a new bottleneck
The AI agent shortened the wait from application to interview. Successful applicants in the remote condition reached the interview after a median of 0.32 days with AI versus 0.51 days with a human.
Human evaluation then slowed the process. The median interview-to-offer interval was 7.24 days after an AI-led interview and 2.62 days after a human-led interview. Total median time from application to job start reached 24 days for the AI group versus 20 days for the human group.
Automation moved the queue downstream. A hiring team can conduct more interviews and still make candidates wait longer if review capacity stays fixed.
The paper's cost model points to the same constraint. AI became cost-effective at lower volume in high-wage settings, while the break-even point reached thousands or tens of thousands of interviews in other scenarios. The business case depends on wage levels, vendor pricing, failure rates, and review work. A demo price per minute cannot answer it.
What did the study not prove?
It did not test autonomous selection
Human recruiters made every offer decision. They knew whether AI or a person had conducted the interview. The experiment therefore supports AI-assisted evidence collection, not an LLM deciding who gets a job.
If a team adds model-generated candidate scores, rankings, rejection recommendations, emotion analysis, or personality inference, it has built a different system with a different risk profile. Any numeric output needs calibration against real outcomes; an LLM confidence score is not a universal probability→.
It did not establish fairness across protected groups
The paper reports that the AI condition did not significantly change the observed gender gap in offers. That result does not establish fairness across race, disability, age, accent, or intersectional groups, and the authors lacked enough pre-treatment detail to identify the origin of all gender differences.
Two controlled studies show why the selection layer needs separate testing. A June 2026 study of Japanese-format resumes held qualifications constant while changing gender-signaling names. Five LLMs produced different scores, and one fairness instruction did not remove the aggregate difference. The study used model-converted resumes and one prompt formulation, so its absolute scores should not be generalized.
A March 2026 Singapore study tested 18 models on synthetic resume variants. Languages, activities, volunteering, and hobbies allowed models to infer demographic attributes after explicit identifiers were removed. The measured disparities belong to that controlled Singaporean setup, but the method exposes a practical failure: deleting names does not delete every demographic proxy.
It did not generalize to every kind of work
The experiment took place at one firm handling high-volume recruitment, primarily for customer-service roles. The authors expect the strongest benefits where interviews repeat, outcomes become observable quickly, and human process variance carries a cost. Specialized roles that depend on tacit knowledge or relationship-building may produce different results.
Applicant surveys also had a 14% completion rate. Surveyed applicants rated the experiences similarly, and 78% of applicants offered a choice selected the AI agent, but those figures should stay tied to this process and applicant population.
A safer boundary for AI job interviews
The study suggests a useful architecture for employers and vendors.

*A defensible workflow separates automated interview delivery from human decision authority, then measures outcomes and provides a route for accommodation or appeal.*
Legal and accessibility checks are part of the system
Rules depend on jurisdiction and on what the tool influences. The EU AI Act lists AI used for recruitment or selection among high-risk employment use cases. Article 6 also contains conditions under which a narrow procedural or preparatory system that does not materially influence a decision may fall outside that classification. Providers claiming that exception must document the assessment. A human in the workflow does not settle the classification by itself.
New York City's Automated Employment Decision Tools rules require a recent bias audit, public audit information, and notices when a covered tool is used. The exact definition and coverage need case-specific review.
The US Department of Justice warns that hiring technologies can screen out qualified people with disabilities. Its AI hiring and disability guidance calls for accessible alternatives, clear accommodation procedures, and tests that measure job skills rather than an applicant's disability.
The NIST AI Risk Management Framework offers a jurisdiction-neutral operating model: govern the system, map its context, measure risk, and manage what the evidence reveals. It is voluntary guidance, not a substitute for employment counsel.
Metrics that prevent a false win
An AI interview pilot needs outcome metrics and process metrics. Track at least:
Compare these measures with a concurrent human process when possible. A shorter scheduling queue can hide a longer review queue. A higher completion rate can hide disparate screening. A cheaper interview can create more expensive appeals.
The useful result is the boundary, not the headline
The field experiment provides strong evidence that a structured AI voice agent can collect useful hiring information at scale. It also shows why teams should resist the broad claim that “AI recruiters outperform humans.” Humans made the decisions, the setting favored repeatable interviews, and the workflow traded faster scheduling for slower review.
Use AI job interviews to standardize a defined evidence-collection task. Keep selection criteria, accommodations, review, appeals, and outcome audits attached to people who can explain and change the process.
FAQ
Did the AI make hiring decisions in the field experiment?
No. The AI voice agent conducted interviews, but human recruiters reviewed interview evidence and made every offer decision. The study does not validate autonomous hiring.
Are AI job interviews legal?
They can be lawful, but employment, discrimination, accessibility, privacy, and automated-decision rules may apply. Coverage depends on jurisdiction and how much the tool influences selection. Employers should review the actual workflow and current local law rather than rely on a vendor label.
Claim checks
| Claim | Status | Evidence boundary |
|---|---|---|
| --- | --- | --- |
| The study followed 70,884 applications and randomized 67,056 eligible applications. | Verified | Voice AI in Firms, experimental sample. |
| AI-led interviews raised the offer rate from 8.70% to 9.73%. | Verified | Randomized human-versus-AI interviewer groups. |
| Job starts and 30-day retention rose by about 18% and 17% on a relative basis. | Verified | Unconditional randomized-sample outcomes. |
| The study found no significant productivity decline. | Qualified | Productivity data covered only a subset of hired workers. |
| The AI workflow took four more median days from application to job start. | Verified | Remote-mode successful hires; 24 versus 20 days. |
| AI interviews are fair across demographic groups. | Rejected | The field experiment did not establish this broad claim; controlled hiring studies find context-specific disparities. |
| A human reviewer automatically removes legal high-risk status. | Rejected | Classification depends on the system's real influence and applicable law. |
| AI interviews always reduce hiring cost. | Rejected | Break-even depended on volume, wages, vendor prices, failures, and review work. |



