ai-labor-market

133 Patent Lawyers, 3 Months of AI: Juniors Gained Most, Kept Least (NBER 2026)

NBER w35720 pre-registered RCT: AI raised patent draft quality 0.38 SD, juniors most. But durable judgment gains without AI went only to seniors (0.45 SD); junior scores split.

作者:编辑兼作者
发布日期:
AI-辅助分析

133 patent lawyers. 11 U.S. IP firms. Three months with an AI drafting assistant. Then a test with the AI switched off. The lawyers who gained the most from AI while using it were the ones who kept the least once it was gone.

That is the headline of NBER Working Paper 35720, released September 1, 2026, by David Autor (MIT) and six co-authors from Google. It is, as far as I can find, the first pre-registered randomized trial that measures both what a professional does with AI and what they can still do without it after months of use. If you draft patents, review them, or supervise the people who do, the numbers below are worth twenty minutes of your attention.

What the experiment actually did

[Fact] Eleven U.S. intellectual property law firms enrolled 156 lawyers; 133 cleared onboarding and were randomized 2:1 into an AI group (90 lawyers) and a no-AI control (43). The tool was InFlow, an unreleased Google Labs writing assistant later folded into Google's other AI products. Assignment was stratified by firm and by experience, with "senior" defined as seven or more years in practice. The pre-registration is public at the AEA RCT Registry (AEARCTR-0015823).

[Fact] There were three graded tasks. At day 10, a patent drafting exercise (two-hour cap, up to 1,000 words). At day 90, a harder drafting exercise (up to 2,000 words) and, separately, a redlining exercise done by everyone without AI: mark up an existing patent application using tracked edits and comments. Every submission was scored on enforceability, accuracy, strategic ambiguity, completeness and clarity by blinded patent attorneys at an independent firm, two raters per piece.

One detail the abstract leaves out: the document lawyers were asked to redline at day 90 was itself an AI-generated draft. So the "judgment without AI" test was, in practice, "can you catch what an AI got wrong."

Finding 1: with AI in hand, everyone drafts better, juniors most

[Fact] AI access raised drafting quality by 0.34 standard deviations at day 10 (p = 0.03) and 0.38 SD at day 90 (p = 0.01). The gain showed up across all five scoring dimensions. It did not come from more excellent drafts; it came from fewer bad ones. The probability of a bottom-quintile score fell by 19 percentage points at day 10 (p = 0.00), and among juniors specifically by 25 points (p = 0.01).

[Fact] AI-assisted lawyers were also a little faster: 10 minutes saved against a 112-minute baseline at day 10 (p = 0.05), and a directionally similar but imprecise 10 minutes against 125 at day 90. Seniors did not speed up at all. In the control group, seniors finished drafts 13 to 24 percent faster than juniors; with AI, juniors matched senior speed and senior quality.

That pattern will feel familiar if you have read the customer-service, writing, or coding studies from 2023 to 2025. AI compresses the gap between novice and expert while the tool is running. This paper reproduces it in a domain where the output is a legal instrument that has to survive a USPTO examiner and, eventually, litigation.

Finding 2: switch the AI off, and the story splits by seniority

[Fact] On the unassisted redlining task, the treated group outscored controls by 0.32 SD (p = 0.04). The entire advantage sat with senior lawyers, who beat their senior controls by 0.45 SD (p = 0.02). The estimated effect for junior lawyers was close to zero, and stayed close to zero whether "junior" was cut at three years, nine years, or split into three tiers.

[Fact] But zero on average hid a lot of movement. Treated juniors produced sharply fewer mediocre (fourth-quintile) scores, more poor (bottom-quintile) scores, and an offsetting rise in good (second-quintile) scores, with no gain at the very top. The authors' own words: the dispersion "is consistent with AI accelerating skill acquisition among some juniors and substituting for it among others," and "our experiment does not answer why."

[Fact] The team also ran a forensic check for lawyers who used AI on the no-AI task anyway. Erring toward over-flagging, they identified 15 of 91 redlining submissions where AI use looked possible. Excluding them costs precision but does not overturn the senior result. A separate re-scoring of every submission by an LLM (Gemini 2.5 Flash) found the same drafting pattern and a smaller, weaker junior improvement on redlining that the human raters did not detect.

Four numbers the paper implies but does not print

[Estimate] Attrition is the first thing I compute in any field trial. Of 156 lawyers recruited, 91 finished the final redlining task: 58 percent. Measured from randomization, 105 of 133 finished day 10 (attrition 21 percent), 98 finished the day-90 draft (26 percent), and 91 finished redlining (32 percent). The authors report no sign that dropout differed by treatment status (p = 0.56, 0.79, 0.63 across the three tasks), which is the check that matters, but a third of the randomized sample is still missing from the headline result.

[Estimate] The junior result rests on a small group. Table 1 puts seniors at 89 of the 133 randomized lawyers (67 percent, average experience 12.4 years), leaving 44 juniors at baseline. The drafting tables report 68 junior rating observations at day 10 and 63 at day 90, which at two raters per lawyer is roughly 34 and 31 juniors respectively. The "juniors bifurcate" finding is built on about three dozen people. That does not make it wrong; it makes it a hypothesis to replicate, not a rule to manage by.

[Estimate] One firm supplied 40 of the 133 randomized lawyers, 30 percent of the sample. Firm fixed effects absorb level differences, but a single-firm culture around AI use can still shape what "junior" means in this data.

[Estimate] Speed versus quality: a 10-minute saving on a 112-minute task is a 9 percent time reduction, bought alongside a 0.34 SD quality gain. If your firm's business case for AI drafting is "faster," this paper says the larger effect is "fewer bad drafts," and the time saving is modest and mostly junior.

The counter-narrative: this is not "AI is deskilling associates"

The headline will get summarized as "AI erodes junior expertise." That is not what the trial shows, and the authors say so. [Claim] Three things separate the actual result from the scary version.

First, the average junior effect is zero, not negative. A negative average would be erosion. Zero with a widening spread means some juniors learned more than their AI-free peers and some learned less. The paper cannot tell you which junior you are.

Second, seniors gained durable judgment from three months of AI drafting. If AI were simply a crutch, the group with the least need for it would show the smallest carryover. They showed the largest. [Claim] The authors' reading is that foundational expertise may be a prerequisite for extracting lasting skill from AI-assisted practice: you need a mental model of a good patent to learn anything from watching a machine draft one.

Third, the redlining task asked lawyers to critique an AI-written draft. [Estimate] Seniors who had spent three months reading InFlow output may simply have learned that specific tool's failure modes, and that is a narrower skill than "professional judgment." The paper's finding survives its own robustness checks; whether it survives a different assistant, or a human-written document, is untested.

The honest summary is not "AI hollows out the pipeline." It is: the productivity gain and the learning gain went to different people, and the profession has no idea yet how to give both to the same person.

What this trial cannot tell you

I want to be direct about the limits, because the design is good enough that the limits are the interesting part.

Six of the seven authors are Google employees, Google paid the experiment's direct costs, the tool was Google's, and every participating firm already drafts patents for Google. NBER papers are circulated for discussion and have not been peer-reviewed. None of that makes the estimates wrong, and the pre-registration constrains the analysis, but the study describes sophisticated firms using an in-house tool at Google's expense, and the results should be read as an upper bound on how well a well-supported AI rollout can go.

There was no baseline skills test. Firms refused it as a condition of joining, so the "learning" inference rests on randomization producing balanced groups, which the balance table supports but cannot prove. The treatment was one tool, one drafting workflow, three months. Whether the junior spread widens into a fan or closes back up over a year is, in the authors' word, "plausible—though not certain" in either direction.

And patent drafting is a specific kind of legal work: technical, rubric-scorable, high-volume. This is not a result about litigation, negotiation, or client counseling.

What it means for your job

If you are a patent attorney or agent, the immediate result is unambiguous: with AI, your bad-draft rate falls, and if you are early-career, your speed and quality converge on your seniors' while the tool is on. That is a real, measured gain, and firms will price it in. Our evergreen analyses of patent attorneys, patent agents and patent examiners already framed drafting as the most exposed task in each role; this paper puts a treatment effect on that exposure.

The harder question is the one the authors leave open. If you are in your first seven years, this trial says AI use will not, on average, cost you judgment, but it may sort you into a group that learned faster or a group that learned less, and you will not find out which until someone hands you a document with the AI turned off. The practical response is to manufacture that test yourself: redline without the assistant at a fixed cadence, and compare.

If you supervise juniors, the senior result is the lever. The lawyers who converted AI use into durable skill were the ones who already knew what good looked like. That argues for pairing AI access with more expert review of AI-assisted work, not less, and for treating the first years of practice as the period when the tool should be most supervised, not least.

For the exposure data behind these roles, see our intellectual property lawyers, patent agents, patent examiners and lawyers pages.

Sources

  • Autor, D., Rodchenko, T., Martin, J., Iscenko, Z., Strand, S., Pearl, D., & Ferere, M. (2026). Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting. NBER Working Paper No. 35720, September 2026. https://www.nber.org/papers/w35720 (DOI 10.3386/w35720). Pre-registration: AEA RCT Registry AEARCTR-0015823.

NBER working papers are circulated for discussion and comment purposes. They have not been peer-reviewed or been subject to the review by the NBER Board of Directors that accompanies official NBER publications. Copyright 2026 by the authors; figures cited here are drawn from the abstract and text with attribution, and no tables or figures are reproduced.

This article was produced with AI-assisted analysis and reviewed for factual accuracy against the source document. Statements tagged [Fact] are taken from the paper, [Estimate] are our own calculations from figures the paper reports, and [Claim] are interpretations, ours or the authors', that the data do not establish on their own.

Analysis based on the Anthropic Economic Index, U.S. Bureau of Labor Statistics, and O*NET occupational data. Learn about our methodology

更新记录

  • 首次发布于 2026年9月7日。
  • 自首次发布以来无实质性更新。

Tags

#NBER#David Autor#patent lawyers#randomized controlled trial#skill acquisition#junior lawyers#AI assistance#expertise

来源

  1. nber.org