ai-labor-market

The Best AI at Doing Your Job Isn't the Best at Helping You Do It

On three of seven real work tasks, a worker model scored better with no AI help at all than with any of the ten AI assistants tested against it. A new UC Berkeley benchmark splits "can AI do your job" from "can AI help you do your job" — and finds the two barely track each other.

PorEditor y autor
Publicado: Última actualización:
Análisis asistido por IARevisado y editado por el autor

On three of seven real work tasks, a worker model scored better with no AI help at all than with any of the ten AI assistants tested against it. Not "worse with a bad assistant" — better alone than with the best available guidance. If you have been reassured that AI won't replace you but someone using AI will, a benchmark released on August 19, 2026 by UC Berkeley's Haas School of Business just made that sentence much harder to say with a straight face.

The paper is called CentaurBench. [Fact] It measures two things that leaderboards have been collapsing into one: can a model do your job, and can a model help you do your job?

Two questions that were being asked as one

Pattaraphon Kenny Wongchamcharoen, Kris Gulati, Min Min Fong, and Abhishek Nagaraj — all at the Data Innovation and AI Lab (DIAL) at Berkeley Haas — built a deliberately asymmetric test. [Fact]

In automation mode, a capable assistant model does the task and hands in the finished output. In augmentation mode, that same capable model doesn't produce the deliverable at all. It writes guidance text for a weaker "worker" model — GPT-3.5-Turbo — which then produces the final output. [Fact]

Why such a weak worker? The authors are explicit: GPT-3.5-Turbo "approximates worker performance on our pilot tasks," and it "represents the cheaper, lower-capability models that are often assigned the execution role in multi-agent settings." [Fact] Hold onto that sentence. It decides how far these results stretch.

Seven tasks, chosen for economic grounding rather than puzzle difficulty: counseling, travel planning, meal planning, operations research analysis, tax preparation, tutoring session planning, and market trends analysis. Six were drawn from GDPval and the Anthropic Economic Index; tax preparation was designed by the authors. [Fact] Ten assistant models were tested, from Claude-Opus-4.8 and Gemini-3.1-Pro down to GPT-3.5-Turbo itself. Outputs were compared blind and pairwise by a panel of LLM judges using task-specific rubrics, replicated across ten runs. [Fact]

The two abilities come apart

Across the nine assistant models appearing in both modes, the rank correlation between automation skill and augmentation skill was ρ = 0.48. [Fact] In five of the seven tasks, the model that won augmentation was a different model than the one that won automation. [Fact] GPT-5-Mini took the top average rank in automation. [Fact]

Here is the part a headline would drop, and it deserves your skepticism: that ρ = 0.48 carries a two-sided p = 0.187 on nine models. [Fact] With a sample that small, the result is not statistically distinguishable from no relationship at all — nor from a fairly strong one. [Estimate] The honest reading is that the paper has shown the two abilities can come apart on specific tasks, not that it has pinned down how loosely they are coupled in general.

Where the help actively hurt

The per-task results are blunter than the correlation. On operations research analysis, tax preparation, and travel planning, the unaided worker model outranked every single assisted condition. [Fact]

Ten assistants, and being left alone beat all of them.

Across all seven tasks, only one assistant — GPT-5-Mini — finished with a better mean rank than the no-guidance baseline, and only barely: 3.66 versus 3.79. [Fact] Nine of the ten assistants made the worker worse on average.

Our own cross-check: the tasks where help hurt are the tasks most at risk

We mapped all seven benchmark tasks onto occupations in our database and pulled our automation risk score for each. The pattern was not the one we expected.

The three tasks where assistance backfired sit at an average automation risk of 48% in our data — tax preparers at 54%, travel agents at 58%, and operations research analysts at 32%. The four tasks where assistance did no net harm average 20.5%dietitians and nutritionists at 12%, tutors at 20%, substance abuse counselors at 20%, and market research analysts at 30%. [Estimate]

That is a 2.3× gap, and it points somewhere uncomfortable: the jobs where AI is most likely to take the whole task from you may also be the jobs where AI is least useful as a partner. [Claim] The middle ground — "I'll just work alongside it" — looks thinnest exactly where people are counting on it most.

Treat that as a hypothesis, not a finding. Seven data points. Our risk scores are our own estimates, not the paper's. And mapping a benchmark task onto an occupation is a judgment call — "meal planning" is one slice of what a dietitian does, not the job. [Estimate]

The obvious objection

The worker in this study is a language model, not a person. People read guidance differently: we ignore advice that smells wrong, we notice when an instruction contradicts a form we have filled out four hundred times, we ask a follow-up question. GPT-3.5-Turbo does none of that. It follows. So the harm measured here is plausibly an upper bound on what happens to a competent human professional. [Estimate]

But that objection cuts less than it looks. The authors picked that worker precisely because it stands in for the cheap execution models now being wired into agent pipelines at real companies. For multi-agent systems this isn't an analogy — it's a direct measurement. [Claim]

One more caveat, and it is the paper's own number turned against it: inter-judge agreement was 71.0% across 6,265 comparisons, but that splits into 74.5% in automation and 67.8% in augmentation. [Fact] The judges agreed less about the augmentation results, which makes the augmentation half of this study the noisier half. [Estimate]

What this changes about how you prepare

Look again at the three tasks where guidance made things worse. Tax preparation, operations research, travel booking. They share a shape: rule-bound work with a verifiable right answer, where a confident-sounding wrong instruction gets executed instead of questioned. In open-ended work — a counseling session plan, a tutoring outline — a bad suggestion is cheap to throw away. In rule-bound work it becomes the output.

So, three concrete things.

Stop reading leaderboards as buying advice. The model that tops automation benchmarks is not necessarily the model that will make you better at your work, and this is the first hard evidence of that gap. Test candidate tools on your own actual task, with your own actual inputs.

Where your work has a checkable right answer, verify against the rule, not against plausibility. The failure mode here is not AI producing obvious nonsense. It is AI producing something that reads correctly and isn't.

Notice the one exception. Tax preparation was the single task where automation ability strongly predicted assistance ability (ρ = 0.85, p = 0.004) — and it was still a task where the unaided worker won. [Fact] Being able to predict which model helps most is not the same as that model helping at all.

None of this says AI assistance doesn't work. It says the assumption that it works — quietly baked into every "AI won't replace you, but…" reassurance — has now been tested on seven real tasks, and it did not come through four of them intact. That is not a reason to avoid these tools. It is a reason to stop taking their benefit on faith and start checking it on the work you actually do.

Sources

  • Wongchamcharoen, P. K., Gulati, K., Fong, M. M., & Nagaraj, A. (2026). CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks. arXiv:2608.18554v1, submitted August 19, 2026. Data Innovation and AI Lab (DIAL), Haas School of Business, UC Berkeley. https://arxiv.org/abs/2608.18554

Automation risk percentages in the cross-check section come from AI Changing Work's own occupation dataset and are not figures from the cited paper.


AI-assisted analysis: this article was drafted with AI assistance and reviewed before publication. Every figure attributed to CentaurBench was verified against the paper's full text; the seven-task risk comparison is our own calculation.

Analysis based on the Anthropic Economic Index, U.S. Bureau of Labor Statistics, and O*NET occupational data. Learn about our methodology

Historial de actualizaciones

  • Publicado por primera vez el 20 de agosto de 2026.
  • Última revisión el 20 de agosto de 2026.

Tags

#arxiv#benchmark#augmentation#automation#uc-berkeley#human-ai-collaboration#ai-labor-market

Fuentes

  1. arxiv.org