ai-labor-market

Brookings: AI Exposure Is the Wrong Yardstick for Workforce Policy

Two MIT FutureTech economists tell Brookings that exposure scores are the wrong basis for policy. The numbers behind their five fixes are thinner than the headline.

作者:编辑兼作者
发布日期:
AI-辅助分析

Roughly 90% of Americans born in the 1940s grew up to out-earn their parents. For those born in the 1980s the figure is 50% and still falling. That is the backdrop against which two MIT FutureTech economists, Guy Ben-Ishai and Neil C. Thompson, published a 29-minute Brookings paper on September 15, 2026, arguing that the number every AI-and-jobs debate starts from, "exposure," is the wrong organizing principle for workforce policy. This post walks through what they found in the economic literature, what they recommend, and where their own evidence is thinner than the headline.

The claim: exposure tells you where AI could act, not where it will

[Fact] The paper's opening move is a rebuke of the metric this site is built on. Exposure scores count the share of an occupation's tasks that a model could in principle perform. Ben-Ishai and Thompson cite Yale Budget Lab work by Gimbel et al. showing that different exposure studies routinely flag computer programmers and telemarketers as highly exposed, yet disagree substantially with each other about which occupations sit at the top.

[Fact] Their four findings from the literature, in their order:

First, technical capability is not commercial automation. Whether a task is actually handed to a model depends on completion time, reliability, error handling, and the cost of wiring it into a workflow.

Second, the question is not jobs created versus destroyed but how AI changes the value of human expertise.

Third, adopting AI well means knowing when to trust it, not just how to prompt it.

Fourth, as models get more reliable at longer tasks, the set of workflows where autonomous use is commercially viable will widen.

None of that is new to readers of Autor and Thompson's expertise framework, which we covered when Brookings floated its all-of-the-above policy menu in July. What is new is the attempt to derive concrete program design from it.

The reliability numbers that carry the argument

[Fact] The paper leans on Mertens et al., the MIT "Crashing Waves vs. Rising Tides" evaluations we summarized in May. As of end-2025, models completed tasks that take a human three to four hours with only a 65% success rate. A tenfold increase in task duration cut model performance by 11%. Healthcare support tasks, at just over an hour, scored highest; legal and architecture-and-engineering tasks, at more than 20 hours, scored lowest.

[Fact] For a human benchmark the authors reach for HEART, a human-reliability framework from industrial safety. Historical human-error probabilities there run from 2% for routine low-skill tasks to 9% for simple tasks done with limited attention and 16% for complex tasks demanding real comprehension.

[Estimate] Put those side by side and a 65% success rate is a 35% failure rate. Against HEART's range that is between 2.2 times (versus the 16% complex-task figure) and 17.5 times (versus the 2% routine figure) the human error rate. The authors themselves flag that the two measures are not directly comparable, and I would go further: HEART was calibrated on nuclear and process-plant operators, not on someone drafting a memo. The direction is still hard to argue with.

[Fact] The paper also says failure rates are halving roughly every 2.5 years, and that if the trend holds models "could achieve a 93% baseline success rate across most professional text-based tasks by 2029."

[Estimate] Those two numbers do not come from the same starting line. A 35% failure rate halving every 2.5 years is about 15% failure at the end of 2028 and about 13% by mid-2029, not 7%. Reaching 93% success by 2029 requires either a higher starting success rate for the "most professional text-based tasks" bucket than the 65% quoted for three-to-four-hour tasks, or a faster halving. The paper does not say which. This is the kind of gap a policy reader should notice before treating 2029 as a planning date.

[Fact] On the cost side, Svanberg et al. (a Thompson-coauthored working paper the authors cite through the Brookings text) find that at today's integration costs U.S. firms would choose to automate only 23% of the computer-vision tasks that are technically automatable.

Expertise, not headcount: the accountant and the family doctor

[Fact] The paper's clearest example pairs two occupations. If AI automates bookkeeping and data entry but stays unreliable on tax advice and analysis, the value of an accountant's expertise rises, wages go up, and employment falls. If AI could reliably diagnose and treat the common conditions a family physician sees, the value of that expertise falls, the work opens up to nurse practitioners, employment rises, and pay drops.

[Fact] A footnote adds the wrinkle that matters most: the same automation could make emergency-room doctors more expert, because they would spend more time on the demanding cases. Expertise effects are occupation-specific, not task-specific.

This is why the authors want policy to distinguish two risks rather than one. Displacement (fewer jobs, higher pay for those left) calls for job-search help, wage insurance, and transitions out. Wage erosion (more jobs, lower pay) calls for advancement pathways inside the occupation.

What they recommend, and what the evidence behind each looks like

[Fact] Five directions, with the numbers the paper attaches to each.

Prioritize actually-displaced workers and high-productivity new opportunities. Davis and von Wachter estimate that displacement in high-unemployment periods costs 2.8 years of prior earnings. Rose and Shem-Tov find low-wage workers still earn 13% less six years after displacement.

Design domain-specific, workplace-based training with employers in the room. Generic AI literacy is explicitly judged insufficient, citing Noy and Zhang and Dell'Acqua et al.

Align programs with the direction of the expertise shift, as above.

Expand apprenticeships. The paper notes that the Department of Labor's newer pay-for-performance grants and the bipartisan LEAP Act tax credit tie public money to verified hires above an employer's baseline, and that earlier federal apprenticeship pushes did not achieve broad employer participation.

Establish a federal wage-insurance program, pointing to Hyman et al.'s finding that the Trade Adjustment Assistance version ran on a self-funding, budget-neutral basis.

[Fact] The authors are candid that the training evidence is weak. They cite Card, Kluve and Weber's meta-analysis and Orrell et al. to say workforce programs "have not produced consistently strong results, particularly when they seek to help displaced workers transition into a new sector or occupation." Their answer is targeting, not abstention.

Where I would push back

The 56-RCT retraining review we covered last week found sector-based programs averaging about +1.7 percentage points in employment per $791 of cost, with the celebrated ten-fold sector effects being headline cases rather than the cluster median. See that post. Ben-Ishai and Thompson's "domain-specific, employer-partnered" prescription is consistent with the sector-program literature, but the paper does not quantify what it expects such programs to return, and the KPMG telemetry we covered in August showed formal AI training effects vanishing within a month. "Domain-specific" is a necessary condition for training to work; the evidence does not yet show it is sufficient.

[Fact] Wage insurance is the least-tested of the five. The paper's support is one program (TAA) and one study. It also proposes targeting "verified AI-related displacement," which requires an administrative test for whether a layoff was AI-caused. The frontier-economy data paper Brookings published a week earlier is, in effect, a list of reasons that attribution is hard today.

[Fact] A disclosure worth knowing: the bibliography's "Ben-Ishai et al." entry on shared prosperity is co-authored with Google's Jeff Dean, James Manyika, Ruth Porat, Hal Varian and Kent Walker. That does not invalidate the argument, but a reader weighing a paper that recommends against "sweeping" interventions such as capital taxes should know the lead author's prior work sits close to a company with a stake in that question.

What this means for you

If you are in an occupation this site scores as highly exposed, the paper's message is that the score is a starting point, not a forecast. The questions that decide your outcome are whether the automatable tasks are the ones that made your expertise valuable, and whether your employer can absorb a 35% error rate on the tasks it hands to a model. For most workers in most firms, right now, the second answer is no.

If you are in a trade or a regulated profession where task durations run to many hours, the duration finding is on your side for now. Our page on welders is a useful contrast with the programmer page linked above.

And if you are a policymaker or an HR lead, the cheapest thing in this paper is also the most actionable: measure which of the two risks, displacement or wage erosion, your occupation actually faces before choosing a program. The evidence behind wage insurance and apprenticeships is thinner than the evidence behind the displacement costs they are meant to offset.

Sources

  • Ben-Ishai, G. and Thompson, N. C., "Workforce policy for the age of AI: Recommendations from the economic literature," Brookings Institution, Economic Studies, September 15, 2026. https://www.brookings.edu/articles/workforce-policy-for-the-age-of-ai/

Figures for Mertens et al., HEART (Williams and Bell), Svanberg et al., Davis and von Wachter, Rose and Shem-Tov, and Hyman et al. are quoted as reported in the Brookings text; underlying papers were not independently re-read for this post. The "35%", "2.2 to 17.5 times", "15%/13% by 2028-29" and "2.8 years" comparisons are this site's arithmetic on the paper's figures.

This article was produced with AI assistance and reviewed before publication. Data tags: [Fact] = reported in the source; [Estimate] = this site's calculation on source figures.

Analysis based on the Anthropic Economic Index, U.S. Bureau of Labor Statistics, and O*NET occupational data. Learn about our methodology

更新记录

  • 首次发布于 2026年9月15日。
  • 自首次发布以来无实质性更新。

Tags

#brookings#workforce-policy#expertise#ai-exposure#apprenticeships#wage-insurance#mit-futuretech

来源

  1. brookings.edu