All posts

From Tokenmaxxing to Token Anxiety

In April, the defining excess of workplace AI was consumption. An employee at Meta built an internal dashboard, reportedly dubbed “Claudeonomics,” that ranked tens of thousands of colleagues by how many AI tokens they burned, awarding titles like “Token Legend” to the heaviest users. The Information’s reporting, widely echoed, described a 22 percent month-over-month jump in internal token consumption, to roughly 74 trillion tokens in a single 30-day window, with no matching gain in output. The practice acquired a name, tokenmaxxing, and a debate: was token consumption a proxy for productivity or a vanity metric? When employees began firing off parallel queries to climb the leaderboard, the question answered itself, and Meta shut the dashboard down.

Four months later the pendulum has swung hard the other way. In June, Bloomberg reported that Uber, having exhausted its annual AI coding budget by April, capped employees at $1,500 per month in token spending per agentic coding tool. Accenture’s agentic AI lead told staff to stop using AI for unnecessary tasks, singling out non-engineers feeding PDFs into models to generate slide decks, and announced an internal system for tracking where token spend goes. The vendors are metering from their side too: Anthropic added weekly rate limits for its heaviest subscribers in August 2025, and Fast Company describes the broader turn as the end of all-you-can-eat AI, with explicit daily caps appearing where vague access language used to stand.

Between the leaderboard and the cap sits the user, and the user is not doing well. Commentators have started calling the resulting condition token anxiety: the hesitation, second-guessing, and defensive under-use that appear once people become conscious of the meter running behind every prompt. One education writer, observing the same condition in students, defines it as “the cognitive overhead of constantly monitoring, budgeting, and optimizing AI consumption, combined with the psychological weight of a tool that never stops being available and never stops implying that you could be more productive.” Students, the author reports, have invented their own coping strategy of rationing their access to the tool, sometimes usefully, sometimes just fearfully.

The last time the meter ran

When I wrote about Legora’s move to consumption pricing in June, I argued that lawyers who trained under transactional Westlaw and Lexis pricing already understood what a metered research tool does: it turns efficient use into a billable skill. Under that older model, a single search could cost on the order of $60 to $100, each out-of-plan document added to the total, and law schools taught cost-effective research as a discipline. I stand by the claim that the discipline was valuable, but it came with a pathology.

The clearest statement of it comes from Beth Wilensky of Michigan Law, in a 2016 article asking when we should teach students to pay attention to the costs of legal research. Her starting point was that we treated cost-consciousness as an unqualified good: of course students should learn how Westlaw and Lexis charge, and of course they should learn to research with the meter in mind. Her conclusion was that for novices, the conventional wisdom is backwards. A first-year student who is taught to watch the meter learns a different lesson than the one intended. She runs fewer searches, opens fewer documents, and stops earlier, not because the research is done but because each further step has a visible price. The floundering that cost-conscious research eliminates is how a novice builds judgment: bad searches teach what good searches look like, and the dead ends teach where the live ones are. Wilensky’s prescription was sequencing. Let students learn to research without worrying about cost, and introduce the meter only once they have the competence to economize without truncating.

The anxiety she was managing was not hypothetical, and it was not limited to students. Law librarians of that era built entire instructional programs around the fear junior researchers felt about running up charges a partner or client would see. And at least some of that fear was miscalibrated. Sarah Gotschall surveyed law students about their summer employment and found that employers placed considerably less weight on containing online research costs than our teaching orthodoxy assumed; the anxiety outran the actual constraint. The pattern is older than Westlaw. Constance Mellon’s foundational work on library anxiety found that most undergraduates described their initial response to research in terms of fear, and that the fear itself, not any lack of ability, was what kept them from doing the research well. Anxious researchers do not become efficient researchers. They become researchers who stop.

What token anxiety costs

Substitute tokens for search charges and the mechanism transfers cleanly, except that the stakes are now higher, because the steps a token-anxious lawyer skips are the steps that make AI output safe to rely on.

The most valuable compute a lawyer spends is the compute spent checking: the grounding-in-quotes pass, the verification loop before finalizing, the corroboration against independent authority that I have argued should be the standard for training new lawyers. Every one of those passes consumes tokens while producing no new work product, which makes them the first thing a lawyer watching a meter cuts. A verification pass on a brief costs a few dollars of compute; a fabricated citation that survives into a filing costs sanctions that ran to at least $145,000 across U.S. courts in the first quarter of this year, and the sanctioned lawyers shared one feature: no functioning verification practice. Token anxiety manufactures that failure.

A lawyer trying to spend fewer tokens will also be tempted to feed the model less: fewer grounding documents, terser instructions, one overloaded prompt instead of a well-scoped sequence. This inverts the practices that produce reliable output. Grounded prompts with the relevant documents attached are what suppress confabulation, and a starved prompt pushes the model back onto its parametric memory, which is where fabricated authority comes from. Skimping on context to save tokens is the equivalent of the associate who, afraid of the per-document charge, cited the headnote without opening the case.

The most anxious response to a meter is not to use the tool inefficiently but not to use it at all, and the education writer observing token anxiety in students reports as much: usage rationed by fear rather than judgment. For lawyers this failure compounds. Competence with these tools is now an ethical expectation under Comment 8 to Model Rule 1.1, clients ask about it in outside-counsel guidelines, and the skill is built the same way research skill is built: by use, including inefficient early use. A firm whose juniors are afraid to run the tool is training a generation that can neither use it nor supervise it.

The token version of Gotschall’s miscalibration is already visible. The anxiety is aimed at the number the lawyer can see, the token count on the dashboard, rather than the numbers that dominate the economics: the lawyer’s own billable time, the client’s exposure, the cost of a defective filing. Harvey’s co-founder has put the compute behind a single document-drafting query at roughly $20. Twenty dollars is four minutes of a mid-level associate’s time. A lawyer who spends thirty minutes hand-rationing prompts to save a few dollars of compute, or who re-does by hand what the tool would have done verifiably well, has not economized. She has transferred the cost to a more expensive meter that nobody put on a dashboard.

Mitigating token anxiety

As lawyers, we have already tested the fixes under the last meter. For a firm rolling out consumption-priced AI, or a lawyer whose employer just installed a spend dashboard:

Sequence the meter after competence. Wilensky’s core insight applies unchanged: let people learn to direct the tool before they learn what it costs. Give trainees a sandbox or an unmetered training allowance, keep per-user dashboards away from them, and introduce cost discipline as a deliberate second course once they can scope, ground, and verify a task competently. The learning-phase tokens are cheap; the habits formed during that phase are not.

Budget ex ante, at the matter level, not prompt by prompt. The worst version of cost control is a thousand individual flinches, each lawyer deciding at each prompt whether this one is worth it. The better version is the one Legora built into its own consumption model: thresholds and alerts set in advance by someone with visibility into the matter’s economics. A lawyer working inside a known envelope can spend freely and well within it. A lawyer with no envelope is back to the thousand flinches.

Make verification non-discretionary. If a firm economizes on AI, the savings must come from scoping and prompt hygiene, the practices that reduce cost and improve quality at the same time, and never from skipping the checking passes. A written norm helps: verification tokens are not overhead to be minimized but the product being purchased. The June post argued that the expensive way to use these tools and the unreliable way are the same way. Token anxiety obscures the inverse: the way that looks cheapest on the dashboard, unverified single passes over starved context, is the most expensive way the moment anything goes wrong.

Publish the actual prices. Anxiety thrives on unpriced dread. A firm that circulates rough per-task figures, a few dollars for a research memo pass, tens of dollars for a substantial draft, alongside the comparison figures that supply perspective (the associate’s hourly cost, the Q1 sanctions numbers), replaces a vague fear with a calibration. Gotschall’s students feared a constraint their employers barely enforced; the cure was information, not reassurance.

Measure output, not tokens, in both directions. Meta’s leaderboard failed because token consumption is not a productivity metric, and a spend dashboard read as a shame index fails for the same reason in reverse. Tokens are an input, like Westlaw charges were: a disbursement to be managed, attributed, and passed through where appropriate. The question to ask of any lawyer’s AI usage is the one that was always the right question about research: did the work product justify the spend, and would more spend have made it more defensible?

The efficiency and quality gains from these tools are available only to lawyers who use them, use them enough to get good, and never skip the checking. Tokenmaxxing wasted money by treating the meter as the point. Token anxiety wastes something more expensive by treating the meter as a threat. Our profession spent two decades learning to do good research under a meter; that hard-won curriculum is sitting in the literature, waiting to be reused.


This post draws on reporting on Meta’s token leaderboard and tokenmaxxing from Forbes, The Pragmatic Engineer, and Built In; Bloomberg and TechCrunch on Uber’s spending caps; IT Pro on Accenture; Fast Company on tightening token limits; and UC Strategies and The Augmented Educator on token anxiety. On research anxiety, it draws on Beth H. Wilensky, When Should We Teach Our Students to Pay Attention to the Costs of Legal Research?, 24 Perspectives: Teaching Legal Research & Writing 41 (2016); Sarah Gotschall, Teaching Cost-Effective Research Skills: Have We Overemphasized Its Importance?, 29 Legal Reference Services Quarterly, no. 2 (2010); and Constance A. Mellon, Library Anxiety: A Grounded Theory and Its Development, 47 College & Research Libraries 160 (1986). It builds on earlier posts on consumption pricing, how lawyers should prompt, the verification standard, and Q1 2026 citation sanctions.