ai-labor-market

AI Training Worked for Exactly One Month. Then Nothing Stuck.

KPMG gave nearly 4,000 back-office workers GenAI for eight months. Their skill never improved — and training's effect vanished within a month. 713,564 prompts say why.

ByEditor & Author
Published: Last updated:
AI-assisted analysisReviewed and edited by author

Formal AI training worked — for exactly one month. [Fact] In a new field study of 713,564 employee prompts at KPMG, workers' prompting sophistication rose in the month they completed formal GenAI training, then dropped back to baseline in the months that followed. Eight months of daily access produced no improvement either. If the plan is that the workforce will simply "pick up AI over time," this is some of the most detailed evidence yet on how that plan is going. Not well.

A rare look inside a firm's actual prompt logs

Most of what we know about workplace AI use comes from surveys — people telling researchers what they think they do. A new working paper by Nicholas Hallman, Zachary Kowaleski, Anu Puvvada, and Jaime Schmidt (arXiv 2608.27364, posted August 27, 2026) works from something much rarer: the raw logs. [Fact] The authors analyzed 713,564 prompts and their LLM responses from 3,925 back-office employees at KPMG LLP, spanning 15 functional areas over eight months of 2025.

[Fact] The telemetry alone corrects the survey record. In 83.9% of employee-months, workers used at least one GenAI tool — where prior survey-based estimates for comparable workers put usage around 32.1%. [Estimate] That is roughly a 2.6x gap between what workers report and what the logs show — our calculation, not the paper's. Whatever else is true about AI at work, self-reports appear to undercount it badly.

The paper's core question is not whether people use AI but how well. The authors define sophistication as "skilled use informed by an understanding of how the tool works and where it can be applied," and measure it three ways: prompt clarity (nine scored dimensions such as specificity, constraints, and acceptance criteria), deliberate strategy use (techniques like role assignment, worked examples, and asking the model to check its own output), and use-case diversity. [Fact] An LLM scored all 158,496 conversations against a calibrated rubric.

Seniority won

The loudest storyline in workplace AI says juniors have the edge: they grew up on these tools, and AI will flatten the value of experience. This dataset says the opposite.

[Fact] All three sophistication measures rise monotonically with rank — staff-level employees score below the sample mean, employees above manager level score above it, on every dimension. The authors attribute this to delegation skill and domain expertise: knowing what to ask an AI for turns out to be much the same skill as knowing what to ask a junior colleague for.

[Claim] This is now a pattern, not a one-off. Anthropic's 400,000-session study of agentic coding, which we covered in August, found the same shape in a completely different setting: experienced users extracted far more from the identical tool, especially when something went wrong. Two independent datasets — one from a professional-services back office, one from software work — both find that expertise compounds with AI rather than being erased by it.

Where sophistication lives — and where it doesn't

Departments were not equal. [Fact] Strategy (+0.30 standard deviations on prompt clarity), Digital Innovation (+0.30), and Project Management led the firm. At the other end sat Accounting & Finance at −0.31 — the only functional area sitting below the sample mean on every usage measure the paper reports.

That last finding lands harder when you notice whose logs these are. [Claim] KPMG is an accounting firm. Even in the back office of a company whose brand is accounting, the accounting function used GenAI least and least skillfully. The authors' explanation is about task structure, not talent: a standalone chat window fits open-ended, iterative work — strategy documents, change plans — and fits poorly into spreadsheet-heavy, tightly proceduralized workflows.

The plateau nobody planned for

Here is the finding that should worry anyone budgeting for "AI enablement."

[Fact] Over the eight-month window, sophistication did not improve. Monthly averages hovered near the sample mean the whole time — no learning curve, no drift upward as employees accumulated hundreds of prompts of practice.

Formal training looked better, briefly. [Fact] In the month an employee completed the firm's GenAI training, all three sophistication measures rose, with effects statistically distinguishable from zero at the 1% level. [Fact] In the months after, the coefficients fall to roughly zero — and one of them, deliberate strategy use, turns negative. The authors' verdict: "any improvement associated with these trainings is temporary."

[Estimate] Read those coefficients side by side — a clear same-month gain, then point estimates at or slightly below zero afterward — and the conclusion is starker than "training fades." The entire measurable benefit is confined to the month of the training itself. That framing is our reading of the numbers, not the authors' phrasing.

Practice didn't work. Training didn't stick.

[Claim] The seniority and department results hint at why: sophistication tracks the work itself — the judgment and delegation habits people bring to it — rather than tool-specific instruction. Which suggests the binding constraint on AI productivity gains is not access, and not even literacy, but changing how people actually work. That is the slowest thing in any organization to change.

What 4,000 people actually do with AI all day

The use-case distribution is its own reality check. [Fact] Writing and communication accounted for 73.3% of conversations. Knowledge retrieval came next at 23.0%, then text-based analysis at 10.7%, software and tool guidance at 9.9%, coding and data analysis at 7.3%, and ideation at just 1.8% (a conversation can span categories, so shares overlap).

[Estimate] Writing outweighs coding roughly 10-to-1 in this back office — our arithmetic. And inside that writing bucket, editing existing text (45.8%) beat drafting new text (34.8%). [Claim] For all the talk of AI as an autonomous author, the modal use inside a real firm looks closer to a very fast copy editor working over human drafts.

What this means if this is your job

For administrative assistants, accountants, and financial analysts, the department results warn against complacency in both directions. Low sophistication in finance functions today does not mean AI is irrelevant to the work — it means the current chat-shaped tools fit the workflow poorly. When AI arrives embedded inside the spreadsheet and the ERP system rather than beside them, that gap can close abruptly.

For project managers and management analysts, the news is nearer-term and better: the open-ended, iterative shape of that work is exactly where sophisticated use already concentrates — and where the seniority premium shows up most clearly.

And for anyone weighing a prompt-engineering course: this study cannot prove a course is worthless — it measured one firm's internal training, not every format. [Claim] But it is direct evidence that a training event without changed daily workflows produces no lasting behavioral change. Whatever "learning AI" is, it looks less like a course and more like a work habit.

The limits worth stating

The design is descriptive, not causal — the authors say so plainly. The data cannot show whether sophisticated prompts produced better output or any productivity gain; sophistication here is a proxy for skilled use, not a measured outcome. It is one firm, one back office, and the eight-month window closed in 2025 — models and prompting norms have moved since, and the authors note their reported behaviors "may already differ" from current practice. [Claim] The seniority gradient could also partly reflect what different ranks are asked to do, not only how skilled they are. None of these caveats touch the headline result, though: within this window, neither time nor training moved sophistication.

Sources

  • Hallman, N. J., Kowaleski, Z. T., Puvvada, A., & Schmidt, J. J. (2026). Sophistication in GenAI Use: Field Evidence from a Large Firm. arXiv:2608.27364. Working paper, not yet peer reviewed. https://arxiv.org/abs/2608.27364 — licensed CC BY-NC-ND 4.0; findings reported here narratively with attribution, no figures or tables reproduced.

AI-assisted analysis: this article was drafted with AI assistance and reviewed by a human editor before publication. All figures were verified against the primary source above.

Analysis based on the Anthropic Economic Index, U.S. Bureau of Labor Statistics, and O*NET occupational data. Learn about our methodology

Update history

  • First published on August 28, 2026.
  • Last reviewed on August 28, 2026.

Tags

#genai adoption#workplace ai#ai training#prompt sophistication#field study

Sources

  1. arxiv.org