A 200-Query AI Mandate Turned 31% of Prompts Into Filler. Halving It Raised Sales 7%
A medical-device firm made 5,000 staff send 200 AI queries a month. 31% were repeats or off-task. Where branches cut the quota to 100, filler made up 90% of the drop and sales rose 7%.
In March 2025, a Chinese medical-device company opened an in-house AI assistant to about 5,000 office staff. By the end of that month, only 14.8% of them had typed a single question into it.
So management did what a growing number of firms are doing: it made AI use mandatory. Every non-production employee had to submit at least 200 queries a month, with a fine for repeat offenders. A working paper posted to arXiv on September 26, 2026 by five economists at the University of Hong Kong tracks what happened next, query by query. The short version is that the quota worked, then it broke, and cutting it in half made the sales team sell more.
The quota got people in the door
[Fact] The mandate did its first job. In April 2025, the month it took effect, 31.2% of employees used the tool for the first time. Another 9.1% joined in May. [Estimate] Adding those up, the share who had ever used it went from 14.8% to about 55% in two months.
[Fact] The penalty was not symbolic. A first miss earned a formal warning; repeated misses cost RMB 200, which the authors put at about 2% of average monthly pay and 7% of average monthly performance pay.
[Fact] The tool itself was practical. It ran on DeepSeek models fine-tuned on the firm's technical files, customer records and regulatory archives. The authors give one example: a salesperson preparing a hospital visit could build an account profile and pitch in about 10 minutes, against 40 to 60 minutes of manual searching.
That is the part of the story that usually gets told. The rest is what the query logs show.
Then the quota became the job
The researchers could read every prompt, not just count them. That is rare. Most adoption studies see logins or seat licenses.
[Fact] Among employees who used the tool at all, 12% of monthly totals landed on exactly 200, and 95% stopped at 220 or below. Almost nobody used the assistant much beyond what was required.
[Fact] Timing gave it away too. In the first half of each month, each workday accounted for roughly 3% of monthly queries. In the final workdays, that jumped to 7% to 8% per day. People were topping up before the deadline.
[Fact] Content was the clearest signal. The authors flagged 31% of queries as either exact repeats of the employee's own earlier prompt that month, or personal and off-task. The published examples of off-task prompts include questions about US tariffs on China and diary-style entries about missing old friends.
The authors cite the term "tokenmaxxing," which the Wall Street Journal used in April 2026 for this behavior in the tech sector. Here it has a measured size.
An accidental experiment
What makes this paper more than an anecdote is how the fix was rolled out.
[Fact] Three months in, branches complained that 200 queries a month was time-consuming. Headquarters lowered the target to 100 in July 2025. The paper notes that headquarters made that call without reviewing usage data or query content. It went out as a CEO-office memo rather than a formal notice, and branches read it differently. 19 branches switched immediately, others later, and 15 kept enforcing 200 through December.
That muddle is useful. It lets the authors compare branches that switched with branches that had not yet switched, month by month. [Fact] Switching and non-switching branches looked alike on headcount, gender mix, age, education and baseline usage (joint test p = 0.49), and event-study plots show flat trends before each switch.
What disappeared when the bar dropped
[Fact] Lowering the target cut monthly queries by about 29 per employee, a 30% drop from the sample mean. Of that decline, roughly 26 queries were repeats or off-task, about 90% of the fall. The change in non-repeated, work-related queries was -2.7, statistically indistinguishable from zero.
[Fact] The share of low-quality queries fell from 31% to 18%, and the end-of-month surge flattened.
Employees who had been sitting right at 200 cut back the most. [Fact] Among those who were not anchored near the threshold, total usage barely moved: low-quality queries fell and high-quality ones rose by about the same amount.
Then the outcome that matters to the firm. [Fact] For the one group with individual performance data, more than 500 salespeople in eight branches, lowering the target raised monthly sales by about 7%. The result survives wild-bootstrap inference and dropping one branch at a time.
A number the authors did not compute
The paper also reports how each type of query relates to sales. [Fact] Each extra repeated or off-task query is associated with 0.025% lower monthly sales; each extra genuine work query with 0.018% higher sales.
[Estimate] Plug in the size of the change. Removing about 26 filler queries is worth roughly +0.65% in sales by that association. The small loss of work queries takes back about 0.05%. Net: about 0.6%, less than a tenth of the 7% effect.
That gap is the interesting part. Most of the sales gain did not come from "better prompts." [Claim] The likelier channel is simpler: salespeople stopped spending the last days of every month feeding a counter and went back to selling. The authors list time freed from quota-filling as one plausible channel. Our arithmetic suggests it is the bigger one. This is our inference from two of their coefficients, not their finding, and the per-query coefficients are associations, not causal estimates.
Where the evidence stops
This is one firm, pseudonymised as "MedTech Corp," in one country, over ten months. Limits worth keeping in view:
- [Fact] Every branch got the 200 mandate on the same day, so the paper cannot say whether the mandate itself was good or bad. It measures only the effect of relaxing an existing one. The authors state this directly.
- [Fact] "Off-task" is a label assigned by an LLM classifier (DeepSeek-V4 Pro), checked against three human annotators. Agreement is high (Fleiss' kappa 0.843 with the model included), but it is still a judgment about intent, not an observed fact.
- [Fact] Sales data cover eight branches. With that few clusters, confidence intervals are wider than the headline suggests.
- [Fact] The logs show what people asked, not whether they used the answers.
There is also a fair counterargument to the whole framing. [Claim] Without the harsh first target, the tool might have stayed at 15% adoption, and the firm would have had nothing to relax. A blunt mandate may be a reasonable way to start. The paper supports that reading too: the case against quotas here is a case against leaving a high one in place after it has done its work.
What this means for your job
If your employer is counting your AI use, this study is a preview.
- A usage count is not a productivity measure. In this firm, 31% of counted activity was filler. Any dashboard that ranks employees by prompts is partly ranking them by patience for busywork.
- If you manage people, read the content before you set the number. Headquarters here cut the quota because of complaints, not data, and still got a 7% sales gain. A firm that looks at the logs can do better than luck.
- If you are subject to a quota, keep evidence of useful work. In this study, the employees who used the tool for real kept doing so after the bar dropped. That is the usage worth being able to show.
For occupation-level data, see our pages on sales representatives, field sales representatives, sales managers and compensation and benefits managers.
Sources
- Jie Gong, Jiayi Hou, Jin Li, Fei Pu and Xinjue Yao, "When Less Is More: Managing AI Adoption with Adaptive Incentive Design," arXiv:2609.32859, September 2026: https://arxiv.org/abs/2609.32859
About this analysis
This article was written with AI assistance. All figures marked [Fact] were read from the full text of arXiv:2609.32859v1 (57 pages, including appendix), not from the abstract alone. The paper is a working paper and has not been peer reviewed. The sales decomposition marked [Estimate] (about 0.6% of the 7% explained by query mix) is our own arithmetic from the paper's reported coefficients and is not a claim the authors make. No tables or figures from the paper are reproduced.
Analysis based on the Anthropic Economic Index, U.S. Bureau of Labor Statistics, and O*NET occupational data. Learn about our methodology
Historial de actualizaciones
- Publicado por primera vez el 30 de septiembre de 2026.
- Sin actualizaciones sustanciales desde la publicación inicial.