Blog
Token-based consumption pricing does not behave like the software line items a chief financial officer knows how to model, and the gap between what engineers consume and what finance expects has stopped being hypothetical. A seat-based SaaS contract is a fixed and easier to forecast the budget because you can count users, do some multiplication, and then you have a budget for the year. A token meter is a variable that moves with how hard your engineers and your agents decide to work on any given day, and at the velocity AI tooling now runs, that variable has started clearing annual budgets in a single quarter.

Uber is the case everyone can see from the outside and been talking about recently. It rolled Claude Code out to its engineers in December 2025, and adoption ran from 32% of engineers in February to 84% classified as agentic users by March, with heavy users spending two thousand dollars a month and the CTO burning twelve hundred in a single two-hour session. The company exhausted its entire 2026 AI budget by April, four months in, and its CTO told The Information they were back to the drawing board on how to forecast any of it. The tool worked. That was the problem. It produced enough value that nobody wanted to slow down, and the meter ran accordingly.
Amazon shows the next failure mode, the one where the incentive itself manufactures waste. Amazon built an internal leaderboard called KiroRank that scored engineers on token consumption, then pulled it after staff began "tokenmaxxing," assigning agents to pointless tasks and spinning loops whose output nobody shipped, purely to climb the rankings. Every fabricated task consumed real compute the company paid for. The board measured tokens burned and got exactly that, which is Goodhart's law with a GPU bill attached: the moment token usage became the target, it stopped being a measure of useful work. Uber had run the same play with its own usage leaderboard, and the same reporting notes Meta employees gaming a similar internal table by driving consumption up for its own sake.
The pattern is not confined to companies large enough to make headlines. In the field I have watched a single engineer at an aerospace firm spend a hundred thousand dollars in Claude credits in four weeks, on a corporate account with no mechanism to notice until the invoice arrived. None of these are anomalies. They are the predictable output of a meter with no governance layer between the engineer and the bill.
The obvious objection is that this solves itself. Frontier models keep getting stronger, and the price per token keeps falling, fast enough that inference for a given task has dropped by orders of magnitude in two years. If tokens get cheaper every quarter, the spend should follow them down. BUT it will not as we know from Jevons Paradox when a resource becomes more efficient, total consumption does not fall, it explodes, because efficiency unleashes demand that was previously priced out. I have argued before that this paradox breaks in B2B outreach, where the demand side is human attention, which is fixed and cannot absorb the new supply. Token consumption is the opposite case because the demand side is uncapped, because an agent will consume exactly as many tokens as you let it. Cheaper tokens do not buy a smaller bill, because you will buy more agent loops, more reasoning steps, more retries, and more context re-read on every call, and total token volume climbs faster than unit price falls. Gartner projects spending on AI agent software will reach roughly two hundred billion dollars in 2026, up more than a hundred and thirty percent in a year. The unit is getting cheaper and the aggregate is running away at the same time, and agent loops generating tokens nobody asked for are the engine of it.
This is where owning your token generation stops being an ideological preference and becomes a financial control. When inference runs on infrastructure you own, the marginal cost of a token collapses toward the hardware you have already amortized and the electricity to run it. The agent that loops a hundred times draws down capacity you have already paid for instead of multiplying a vendor invoice by a hundred. Your worst case becomes saturation of your own hardware, a utilization problem you can watch and schedule against, rather than a budget overrun you find out about after it has happened. That is the relief valve, and it is the insurance. It does not make tokens free, but it converts an uncapped variable cost into a fixed cost with a ceiling you set, and it puts the kill-switch on a runaway loop inside your own control plane instead of behind a vendor's API. Western open-weight models now reach frontier quality, so the choice no longer trades capability for control.
The honest cost is talent. Running your own generation demands FinOps discipline fused with AI infrastructure expertise: capacity planning, utilization economics, per-team token observability, quotas, and guardrails that cap autonomous loops before they spend anything. That skill set is scarce relative to how fast demand for it is growing, and it is the real reason cloud-first consumption remains the right call when you are building from zero to one or running genuinely unpredictable workloads. You should not stand up a GPU fleet to run an experiment. But once consumption is your dominant financial risk rather than your smallest line item, the scarce hire is cheaper than the overrun, and the relief valve pays for itself the first quarter an agent loop would otherwise have detonated the budget.
The decision rule is a single question every operator at scale should be able to answer, and most cannot. If your agent usage multiplies next quarter, and Jevons says it will, is your token bill bounded by hardware you own or unbounded by a meter you do not control? Owning your token generation is how you make that number bounded before the volume arrives instead of after. The relief valve is not about saving money at today's consumption. It is insurance against the consumption that cheaper, stronger models are about to make inevitable, and the only real question left is whether you own the thing the meter is counting.
