ai-labor-market

56 Randomized Trials Say Retraining Works, Barely. Here Is What Does

$13,598 per person, +1.7 points of employment, +$800 a year. The largest meta-analysis of U.S. job training yet, and the one program type that beats it tenfold.

লেখক:সম্পাদক ও লেখক
প্রকাশিত:
AI-সহায়ক বিশ্লেষণ

$13,598 per person. Employment up 1.7 percentage points. Earnings up about $800 a year. That is what the average U.S. job training program delivers, averaged across every randomized trial run since 1973. If "retraining" is your plan for surviving AI, those three numbers are the plan's track record.

The source is a 125-page evidence review by David Roodman and Maxim Massenkoff, published by Anthropic on August 12, 2026 as Anthropic Institute Working Paper 2026-01 and posted to arXiv on September 7 (2609.07011). It is, by the authors' count, the most exhaustive meta-analysis of randomized job training experiments ever assembled for the United States: 146 impact estimates from 56 trials. It also finds one cluster of programs that beats the average by an order of magnitude, and explains why that cluster is so hard to copy.

We read the full PDF, not the abstract. Below is what it says, what it does not say, and what it means if you are looking at your own job and wondering whether "just retrain" is a real answer.

Who wrote this and why it matters for how you read it

[Fact] Roodman is an independent researcher; Massenkoff is at Anthropic. The paper's opening sentence is about AI: "Among experts and the public, worker retraining is the most popular policy for mitigating labor market disruption from AI." That popularity claim rests on a survey of economists, AI experts and superforecasters (Karger et al., NBER Working Paper 35046, which we covered here), in which "retraining support" beat enhanced unemployment insurance, a jobs guarantee and universal basic income as the preferred response to AI-driven job loss.

[Fact] The extraction of program traits and impact numbers from thousands of pages of government evaluation reports was done with Claude Opus, under what the authors call "engaged oversight," including an adversarial second pass in which the model re-read its own reasoning table against the source documents.

Two disclosures of our own. This site's occupation-level exposure data is itself built in part on Anthropic's Economic Index releases, so we are not a neutral party when an Anthropic paper lands. And the paper is a working paper: it has not been peer-reviewed. We checked every number below against the PDF, but the numbers are theirs, not an independent replication.

The average program: real, small, and expensive

[Fact] The 56 trials cover programs that mostly served low-income adults and youth, training them for jobs such as nursing aide, IT support technician and welder. The typical program ran about six months. Among the 67% of "training-primary" programs that reported costs, the average was $13,598 per person offered a slot, in 2025 dollars.

[Fact] Pooling the trials with random-effects meta-analysis, the programs raised employment by 2.8 percentage points in year two and 1.7 points in years three to five, against a control-group employment rate of about 63%. Pre-tax earnings rose $1,139 a year in year two and $791 a year in years three to five. All four averages are statistically significant. The effects are per person offered training, not per person trained.

That last clause carries more weight than it looks. [Fact] In these trials, people offered a program were 66 percentage points more likely to take that specific program than the control group, but only 28 points more likely to receive any training, because control-group members went and found training elsewhere. The authors estimate that the effect per actual trainee is roughly two to four times the headline number, and call that an upper bound.

[Estimate] Here is one way to feel the size of $791 a year. Dividing the $13,598 average cost by the $791 long-run annual gain gives 17.2 years of undiscounted earnings gains to pay back the training cost. Most of the trials stopped following people by year six. So whether these programs "break even" depends almost entirely on an assumption the data cannot check: that a gain measured at year five persists for another decade or more. The authors say so themselves, calling their break-even assumptions "reasonable but debatable."

[Fact] Under those assumptions, the net present value of earnings gains to retirement is $14,146 per person, rising to $19,525 with fringe benefits, against the $13,598 cost. The narrow benefit-cost ratio is 1.44. From the government's side, higher taxes, lower benefit use and reduced spending on other training claw back about three-quarters of the outlay, leaving a net fiscal cost of 24 cents per dollar spent. The authors' summary phrase is "a little better than break-even."

The big federal programs did no better, and sometimes nothing

If the average program is modest, the programs run at national scale, the scale you would need to buffer an AI shock, are the ones to watch. [Fact] The U.S. has randomized evaluations of three of them:

The Job Training Partnership Act, evaluated in the late 1980s, lifted adult employment about 2.3 points from a base near 70% and annual pay by about $1,100 in 2025 dollars. It showed no clear benefit for youth.

Job Corps, the intensive residential youth program dating to the 1960s war on poverty, showed no effect on employment or earnings in follow-ups running 20 years, apart from a transitory one-to-two-point employment bump in years three to five.

The Workforce Investment Act, JTPA's successor, returned "essentially zeros" for both low-income adults and dislocated workers in the WIA Gold Standard evaluation. One reason: the treatment-control gap in actually receiving training was just 15 percentage points, because many people offered training declined and many controls found it anyway.

[Estimate] The paper cites a GAO figure that in 2017 the federal government spent $14 billion a year on job search assistance, counseling and retraining, reaching an estimated 10.7 million people. That works out to about $1,300 per person, roughly one-tenth of the $13,598 per person in the randomized trials. The comparison is imperfect, since the GAO count includes light-touch job search help, but it suggests the federal system mostly does not deliver "training" in the sense the trials measured. Scaling the trial results to the federal budget line is not straightforward in either direction.

The exception: sector programs

[Fact] One cluster stands apart. "Sector programs" screen applicants intensively, design curricula with employers, track local skill demand, and place graduates into specific high-demand industries. The best known are Project QUEST in San Antonio, Per Scholas (IT), and Year Up. The paper's forest plots for the broader family of 19 randomized sector-program evaluations show a partially transient employment bump of about 3.0 points and a more durable earnings gain of $3,000 to $4,000 a year in 2025 dollars, with the long-term (three-to-five-year) estimate at $3,703.

[Fact] Costs were not higher. Sector programs averaged $11,602 per person, below the $13,598 of the general pool. The benefit-cost ratio comes out at 7.18 against 1.44, and from the government's perspective the programs pay for themselves twice over through taxes and benefit savings. The societal internal rate of return is 34% versus 6%.

[Fact] The star performers do even better. Year Up's treatment group, after a year of lower earnings during training and internship, pulls ahead by about $2,000 per quarter and stays there, roughly $8,000 a year. The abstract's "ten times" figure refers to programs in this tier, which lifted pay by $5,000 to $10,000 a year.

[Estimate] Worth stating plainly: the "ten times" is the headline, not the family average. Comparing the long-term earnings effect for the sector family ($3,703) with the training-primary pool ($791) gives a ratio of 4.7, and the paper's own text for the broadened family says "3-5 times higher." Ten times is what the best three or four programs achieve. Both numbers are true; they describe different things.

Why the exception is hard to copy

This is the part of the paper that should temper anyone's plan to "scale up sector programs" as an AI response.

[Fact] Selectivity. The four WorkAdvance sites accepted about 20% of applicants after "intensive screening" for hard-to-measure traits like commitment. Per Scholas received 70,000 applications for 5,000 spots in 2025.

[Estimate] That Per Scholas figure is a 7.1% acceptance rate. Roughly 93 out of every 100 people who wanted in were turned away. The executive summary's phrase is that sector programs "filter out >80% of applicants"; at Per Scholas the filter is closer to 93%.

[Fact] The people who get in are already better positioned. Long-term employment among control groups in sector-program trials was 80%, versus 60% in the general training pool, and sector-program controls earned 50 to 75% more. These programs draw from a population that would do relatively well without them.

[Fact] Replication is the graveyard. When the Center for Employment and Training model was replicated at 14 sites, evaluators judged only six achieved high fidelity on employer involvement, and only four replicated the model with high overall fidelity. Among the four PACE-evaluated programs that best fit the sectoral label, one (Year Up) worked and three "achieve no such success." Year Up itself, the authors note from a conversation with its former chief research officer, expanded to new cities rather than scaling in Boston because each city has only so many "deep-pocketed corporate partners" willing to fund training and commit to hiring.

[Fact] Time limits. Effective sector programs teach skills learnable in months. Year Up runs a year, half training and half internship, and that is long by sector-program standards.

What the paper says about AI, and what it does not

We want to be exact about attribution here, because the paper's AI framing is its own but the occupation-level mapping is ours.

[Fact] The authors are direct: "We have no research evidence on how well job training helps people adjust when large language models start doing their jobs." What they offer is extrapolation from the history above, and it points in two directions at once.

[Fact] For lower-wage exposed roles, the paper is cautiously hopeful. It reasons that sector programs teaching certifications in under a year "could make a difference for many workers," and that if demand for training surges, take-up rises and the realized effect could exceed the trial averages, since the per-trainee effect is roughly twice the per-offer effect. It cites Manning and Aguirre's mapping of 356 professions in which 1.7 million "secretaries and administrative assistants" score high on AI exposure (59%) and low on adaptive capacity (14%), alongside payroll clerks and general office clerks.

[Fact] For higher-skill professionals, the paper is blunt: short programs "are ill-matched to the needs of more-skilled workers." Mass displacement of accountants, lawyers and coders would be "unprecedented," deep reskilling "takes years," and a years-long pipeline breaks the sector model because no intermediary can predict which occupations will be in demand four years out. The one piece of evidence on deep retraining it cites, a Danish study of injured tradespeople, found those who went on to higher education earned 25% more than before, but only after a typical four years of college.

[Claim] Our interpretation, not the paper's: this splits the AI question along a line our occupation pages already draw. If you work in administrative support, as an office clerk, a payroll clerk, or in data entry, the retraining evidence is about people like you, and it says a good sector program into IT support or a skilled trade such as welding has a measured, repeatable, if selective, track record. If you are an accountant, a lawyer, a paralegal or a software developer, the evidence base has almost nothing to say about you, and the authors' best guess is that your version of retraining looks like "a compressed mid-career master's degree," not a six-month program. Even truck drivers, whom the cited exposure index rates as barely exposed, get a warning: broader automation could hit that role regardless of how "AI" is defined.

We wrote about the limits of retraining as a policy in March, leaning on commentary. This paper is the quantitative floor under that commentary.

The policy asks, in the authors' words

[Fact] The paper closes with four recommendations: short-duration sector programs could help some lower-skill displaced workers; they will not meet the needs of more-skilled ones; the Trade Adjustment Assistance model, which extends unemployment insurance from six months to three years while funding reskilling, "is worth considering" for people demonstrably displaced by AI, with the caveat that trigger-based aid does not fire when firms simply stop hiring rather than firing; and funders should run a "fire drill," rapidly scaling and randomizing a leading sector program to learn which components are essential before the need is acute.

[Claim] The authors' bottom line, which we share: current programs "would not leave a dent in a persistently high unemployment rate," and the right response is "a combination of partial responses" developed in parallel. Whether or not AI produces mass unemployment, "we will not regret having learned how to best help workers adapt."

What this review cannot tell you

The limits, stated by the authors and by us:

The evidence is American. The paper reviews seven European randomized or discontinuity studies and finds them uninformative on training effects. Korean readers should note that nothing here evaluates Korea's own vocational training system; the transferability of U.S. results to a labor market with different hiring norms and employer structures is an open question, not a finding.

It covers stand-alone programs only. Apprenticeships, community colleges, degree programs and online education, which may matter more for the "compressed master's degree" path, are out of scope.

The populations were mostly low-income or persistently jobless. A wave of displaced accountants would be, in the authors' words, "the most educated and highly paid in history" among the long-term unemployed, with more savings and better job-search skills, and the trials do not describe them.

And the study's own tools are the subject of the study. A paper about whether people can adapt to AI, written with AI, funded by an AI company, is not disqualified by that fact, but readers should hold it in mind, as should we.

If this is your job

Three things follow, none of them "panic."

If your role sits in the exposed-clerical band, look for programs that name the employers who hire their graduates, publish placement rates, and turn people away. Selectivity is not a bug in the evidence; it is the single most consistent feature of the programs that worked. Ask any program how many applicants it accepts.

If your role is professional, treat the retraining question as a multi-year education question and plan its financing accordingly. The one durable data point on deep reskilling is four years of school for a 25% raise.

Wherever you sit, the number that predicts the most in this literature is not the curriculum. It is whether an employer has committed, in advance, to hire. Your occupation page on this site links to the current exposure data for your role; the retraining evidence above is the other half of the picture.

Sources

  • Roodman, D., & Massenkoff, M. (2026). An evidence review of worker retraining. The Anthropic Institute, Working Paper 2026-01, published August 12, 2026. https://www.anthropic.com/research/reviewing-the-evidence-on-worker-retraining-programs. arXiv:2609.07011 (submitted September 7, 2026), https://arxiv.org/abs/2609.07011. Code and data: github.com/droodman/job-training-meta-analysis.
  • Karger, E., et al. (2026). Forecasting the economic effects of AI. NBER Working Paper 35046 (cited by the review for the policy-preference survey).

The arXiv version is distributed under a CC BY-NC-SA 4.0 license. Quotations here are short and attributed; no tables or figures are reproduced. The paper is a working paper and has not been peer-reviewed.

This article was produced with AI-assisted analysis and reviewed for factual accuracy against the source PDF. Statements tagged [Fact] are taken from the paper, [Estimate] are our own calculations from figures the paper reports, and [Claim] are interpretations, ours or the authors', that the data do not establish on their own.

Analysis based on the Anthropic Economic Index, U.S. Bureau of Labor Statistics, and O*NET occupational data. Learn about our methodology

আপডেট ইতিহাস

  • ৯ সেপ্টেম্বর, ২০২৬ তারিখে প্রথম প্রকাশিত।
  • প্রথম প্রকাশের পর থেকে কোনো উল্লেখযোগ্য আপডেট নেই।

Tags

#Anthropic#worker retraining#job training#meta-analysis#sector programs#randomized controlled trial#Year Up#Per Scholas#AI displacement#workforce policy

সূত্র

  1. anthropic.com
  2. arxiv.org