Part one ended on 2 numbers & 2 clocks. A daily allowance that expires at midnight, a monthly one that thins out by week 4.
Those numbers kept me wondering. So I went looking for examples when a token meter has no obvious ceiling.
Harvey, the legal AI company, used 1 trillion tokens in January. By May it was on pace for 12 to 13 trillion a month. Twelvefold in 4 months. Their CEO Winston Weinberg described what is coming in terms lawyers understand. “When you get a bill from a law firm, it says in six-minute increments what people did, and then the hourly rate. Why did they do that? It’s because they’re trying to show ROI.”
A little later, he put the real question on the table: “I just spent $1 billion on tokens. Where’s my ROI?”
Hebbia, which does research work for investors & bankers, processes more than 250 billion tokens a month. More revealing than the number is something they published about where those tokens go. In production, input tokens can outnumber a user’s actual question by ten to fifty times. System prompts, conversation history, retrieved context, function definitions.
The question can be a rounding error. Most of the meter is everything required to make the question intelligible. Seen this way, token spend starts to look less like the price of intelligence than the carrying cost of context.
An agent reading a contract is not expensive b/c it is thinking hard. It is expensive b/c it may read the contract again on each turn. Then the last nine attempts. Multi-agent systems make this stranger still: each participant inherits pieces of what came before, then passes much of it forward again.
Context accumulates. Then context gets sent again. A surprising part of the meter is not new thought at all. It is memory, re-sent.
Which makes for an odd bill. The answer may be the smallest part of what was purchased. Much of the cost lies in reconstructing, again & again, the conditions that made the answer possible.
Another Weinberg line brings the argument back to earth: “You don’t want frontier intelligence running every task. It’s too expensive.”
A change-of-control review may justify it. A first-pass summary probably does not. That sounds like a model-routing decision. It’s really a judgment about where attention may need to go. The same judgment made by a good editor deciding which paragraph needs another hour, or a good partner deciding which document must be read b4 a meeting.
It keeps coming back to one word: attention. If so much of the meter is context rather than conclusion, perhaps that is what is being bought.
When a firm bills in six-minute increments, it is pricing attention rather than insight. If a machine now bills much the same way, has anything actually changed about what was always being sold?
Consider a single change-of-control clause found on page 400 of a deal room. Its value exists b/c pages 1 through 399 were read and ruled out. So which pages were the work?