In November 2021, a million tokens of GPT-3-class output cost US$60. By late 2024 the same capability sold for six cents — a 1,000-fold decline in three years, roughly 10x every year. Andreessen Horowitz's Guido Appenzeller, who measured the fall by matching models on MMLU benchmark scores, named the rate 'LLMflation'. Nothing in the history of computing has deflated faster: not compute in the PC era, not bandwidth in the dotcom buildout.
The bills went the other way. Google processed 480 trillion tokens a month across its surfaces in May 2025, fifty times the year before. By November the figure was 1.3 quadrillion. Goldman Sachs Research now forecasts industry-wide consumption of 120 quadrillion tokens a month by 2030 — a 24-fold multiple of 2026 levels. Enterprise AI invoices routinely land above forecast.
This tension has a name that predates the transistor by eight decades. William Stanley Jevons observed in 1865 that more efficient steam engines increased Britain's coal consumption rather than reducing it, because efficiency made steam power economic in places it had never reached. Tokens — the metered unit in which every large language model transaction is priced — are repeating the pattern at data-centre scale.
The Meter Behind Every Model Call
Every interaction with a large language model is billed in tokens, and the meter runs in both directions. Models from OpenAI, Anthropic and Google split text — increasingly images and other data too — into small units, process them to interpret a request, and generate more of them as output. Pricing is quoted per million tokens, input and output separately. Output carries the premium: five to six times the input rate on published frontier list prices.
Deloitte's advice to boards frames the accounting shift plainly: tokens are a consumption-metered cost centre, closer to cloud compute or electricity than to the software licences most IT budgets were built around. Spend moves with workload mix, model choice and prompt design rather than headcount. Traditional forecasting handles it badly.
Where the volume comes from is a ladder. Inference is the everyday floor: a question in, an answer out, tokens burned on both sides. Analysis scales it up — contracts, filings, customer records and market intelligence fed through models hundreds of pages at a time, input-heavy and continuous. Reasoning multiplies it. Modern models work through problems in internal chain-of-thought steps, weighing options, verifying logic and calling external tools — databases, code executors, web search — along the way. Those intermediate 'thinking' tokens are billed like any others. The user never sees most of them.
Why Are Enterprise AI Bills Rising While Token Prices Collapse?
Falling unit prices expand the set of economically viable workloads faster than they shrink the cost of existing ones — the Jevons paradox, working as written. The Coal Question described efficient engines making steam power viable in mills, ships and railways that could never have justified it before. Substitute inference for steam and the mechanics transfer intact.
The recent data reads like a controlled experiment. On OpenRouter, a marketplace that routes traffic across hundreds of models, the average price paid per million tokens fell to roughly a third of its February level — while tokens consumed quintupled. Weekly volume through the platform grew from about 300 billion tokens to just under 6 trillion in a year, a 19x increase. Ramp's enterprise spending data shows the average price paid collapsing from around US$10 to US$2.50 per million tokens over a single year. Total spending still rose.
Cheaper intelligence keeps unlocking work that was uneconomic at the previous price. Context windows now hold entire codebases. Monitoring that once ran on schedules runs continuously. A one-shot chat gives way to a reasoning loop that consumes far more for the same business outcome — industry estimates put the multiplier anywhere from 5x to 30x. Epoch AI adds the supply-side view: the inference price required to hit a fixed performance milestone has fallen between 9x and 900x per year, depending on the task.
One caveat sits in the data. Google's token growth is decelerating in relative terms — volumes doubled from one quarter to the next, then grew 33% in the following one. The percentages are easing even as the absolute additions set records. Jevons curves do not climb forever; they climb until the new applications saturate. No one yet knows where that is.
Where the Tokens Actually Go
The marginal token is no longer typed by a person; it is generated by machines conferring with themselves. Agentic systems take a single business goal and decompose it into a loop — plan, act, check the result, refine — that can trigger ten to twenty model calls where one used to suffice. A simple chat exchange costs a few thousand tokens. An agent handling multi-step research, coding or process automation consumes tens or hundreds of thousands. Sometimes more. Four workload families dominate current enterprise deployment:
- Coding assistants that debug, refactor and build software, with whole repositories held in context.
- Customer-service agents that resolve complex, multi-turn inquiries end to end.
- Research and strategy tools that synthesise market data into positioning and risk decisions.
- Document-heavy workflows in legal, finance and operations, where models read far more than they write.
Pricing is bifurcating underneath this demand. Commodity inference — classification, extraction, routine summarisation — is heading toward zero, served by open-weight models such as DeepSeek at cents per million tokens. Frontier inference is holding firm or rising, because reasoning-heavy agent workloads need the strongest models and tolerate the premium. July 2026 list prices make the spread explicit.
| Tier | Representative models (July 2026 list) | Input / 1M tokens | Output / 1M tokens | Typical duty |
|---|---|---|---|---|
| Frontier | Claude Opus 4.8 · OpenAI GPT-5.6 · Gemini 3.1 Pro | US$2–5 | US$12–30 | Multi-step reasoning, agent orchestration, high-stakes analysis |
| Mid-tier | Claude Sonnet 4.6 · Gemini 3.5 Flash | US$1.50–3 | US$9–15 | Production workloads balancing quality against cost |
| Commodity / open-weight | Claude Haiku 4.5 · DeepSeek V4 · Gemini 2.5 Flash-Lite | US$0.10–1 | US$0.28–5 | Classification, extraction, routing and summarisation at volume |
List prices are the ceiling, not the invoice. Cached-input rates, batch-processing discounts and prompt-level engineering move real unit economics substantially — which is why token cost management is hardening into an operational discipline of its own.
How Should the C-Suite Govern Token Spend?
Govern tokens the way mature organisations govern cloud and energy: metered, forecast, owned by someone, and engineered down — never renewed annually and forgotten. Deloitte's framing is FinOps applied to AI: visibility into workload demand, governance that ties consumption to business outcomes, and infrastructure choices grounded in unit economics rather than vendor preference.
Three consumption models now compete for enterprise workloads, and each relocates risk rather than removing it. Packaged AI features inside existing software carry predictable fees but hide token efficiency entirely — a meter you cannot see is a meter you cannot optimise. Direct API consumption offers full transparency and full price volatility with it. Self-hosted 'AI factories' demand heavy capital expenditure and scarce talent, and repay both with superior long-run unit economics on high-volume, predictable loads. The pattern emerging in practice is hybrid: APIs for flexibility, dedicated infrastructure where volume justifies it, and open-weight models held as negotiating leverage against lock-in.
The optimisation levers are known and mostly unglamorous. Model routing sends simple tasks to cheap models and reserves premium reasoning for problems that earn it. Prompt and context caching cut the cost of repeated input. Batch APIs trade latency for discounts. Quantisation and early-exit reasoning trim compute per answer. None of this is exotic; all of it compounds. Organisations that treat tokens as an unmanaged variable cost overrun their budgets. Those that engineer for productivity per token build a durable cost advantage. The near-term agenda is concrete:
- Meter first. Measure token spend against business outcomes today, not after the first budget overrun.
- Build 'tokenomics' literacy across finance and engineering, the way FinOps grew up around cloud.
- Pilot agents under hard budget caps — agentic workloads are where consumption multiplies.
- Plan for continuous usage. Vendor and infrastructure strategy should assume AI that runs always-on, not on demand.
Tokens as Strategic Infrastructure
The supply side is being financed as though the demand forecasts were conservative. The five largest US infrastructure providers have committed roughly US$660–690 billion in capital expenditure for 2026, nearly double the 2025 level. Microsoft has disclosed an US$80 billion backlog of cloud orders it cannot fulfil for want of power. Goldman Sachs analyst Jim Schneider expects chip supply to remain short for the next twelve months.
The demand curves justify the concrete. Goldman Sachs Research projects daily model queries reaching 11 billion by 2030 — a 40% compound annual growth rate — with semiconductor providers cutting inference cost per token by 60–70% a year underneath it. Gartner forecasts a 90% reduction in frontier-model inference costs by 2030. Epoch AI projects that training will hold at roughly 5% of total AI compute; the rest is serving models already built. The economics of AI are becoming the economics of running it.
The pattern rhymes with earlier utilities. Electricity reorganised factories; computing reorganised offices; metered intelligence is positioned to do the same to knowledge work — new business models, one-person companies with the output of departments, automation reaching work that never saw software. The winners will not be the organisations that consume the most tokens. They will be the ones that extract the most value per token, on a budget someone actually owns.
The same Goldman Sachs forecast that sees 120 quadrillion tokens a month by 2030 contains a quieter number: by that year, just 12% of knowledge workers are expected to be using the agentic systems consuming them. The infrastructure is arriving a decade ahead of its workforce.
Frequently asked questions
What is a token in LLM pricing?
A token is the unit large language models use to break down and generate text — roughly three-quarters of an English word. Providers bill per million tokens, with separate rates for input (what the model reads) and output (what it writes). On frontier models in mid-2026, output costs five to six times more than input. Reasoning models also bill the internal 'thinking' tokens they generate while working through a problem, even though those tokens never appear in the response.
What is the Jevons paradox in AI?
The Jevons paradox, from William Stanley Jevons's 1865 study of coal, holds that making a resource cheaper to use tends to increase total consumption rather than reduce it. Applied to AI: the price of a fixed level of model capability has fallen roughly tenfold per year since 2021, yet consumption is climbing far faster — Google alone went from 480 trillion to 1.3 quadrillion tokens a month during 2025 — so aggregate enterprise AI spending rises even as unit prices collapse.
How much do frontier LLM APIs cost in 2026?
Published list prices in July 2026 put frontier models at roughly US$2–5 per million input tokens and US$12–30 per million output tokens: Claude Opus 4.8 at US$5 in and US$25 out, OpenAI's flagship GPT-5.6 tier at US$5 in and US$30 out, Gemini 3.1 Pro at US$2 in and US$12 out. Mid-tier models run about US$1.50–3 in and US$9–15 out, and open-weight options such as DeepSeek serve high-volume work for under a dollar per million tokens. Cached-input rates and batch discounts reduce effective prices further.
Why do AI agents consume so many more tokens than chatbots?
An agent decomposes one goal into a loop of planning, acting, checking and retrying, triggering 10 to 20 model calls where a chatbot makes one. Each step carries context, reasoning tokens and tool outputs, all of which are billed. A simple chat exchange uses a few thousand tokens; an agentic workflow for research, coding or process automation consumes tens to hundreds of thousands, with industry estimates putting the multiplier at 5 to 30 times for the same business outcome.
Will enterprise AI costs keep rising if token prices are falling?
Unit prices should keep falling — Gartner forecasts a 90% reduction in frontier inference costs by 2030 — but volume is forecast to grow faster: Goldman Sachs projects a 24-fold increase in monthly token consumption between 2026 and 2030. Total spend therefore depends on governance. Organisations that route tasks to appropriately sized models, cache repeated context and cap agentic pilots can hold costs flat; those that treat tokens as an unmanaged variable cost tend to see bills exceed forecasts.
Sources and further reading
- Andreessen Horowitz — Welcome to LLMflation: LLM inference cost is going down fast (Guido Appenzeller)
- Andreessen Horowitz — Jevons or Bust: token consumption and pricing data
- Goldman Sachs Research — AI agents forecast to boost tech cash flow as usage soars
- Deloitte — How to navigate the economics of AI
- The AI Insider — How AI token supply shapes prices and investor returns
- DevTk.AI — AI API pricing comparison, July 2026
- Epoch AI projections — inference compute to dominate by 2030 (via Crypto Briefing)
Related resources
Go deeper on this topic
Reader notes
Questions, corrections, and field notes
Curated notes from verified readers. Submissions are reviewed before publication.
Loading reader notes...


