According to BBC News, Microsoft, Google and Anthropic have invested hundreds of billions of dollars building the Large Language Models (LLMs) that power services such as ChatGPT, Claude and Gemini. Yet the companies selling those services — and the third-party vendors building on top of them — are finding that setting a price is surprisingly difficult.
The Pricing Puzzle
"Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn't make any sense, honestly, because we don't know," Simon Gooch of Saviynt, an identity management company incorporating agentic AI into its services, told the BBC. The core difficulty lies in tokens — the mathematical building blocks into which user prompts are broken down before being processed by an LLM, and into which responses are converted back into text or commands.
The process is not predictable. BBC reported that subtle variations in a prompt can produce different answers, the same prompt will not always produce the same answer, and different models will produce different answers. Agentic systems — where businesses run multiple AI agents together to make decisions — increase token use and unpredictability further.
Token Economics Shift
Although the cost of individual tokens has plummeted in recent years, according to analysis by Goldman Sachs reported by BBC, the number of tokens consumed has "skyrocketed." The bank forecasts token consumption will increase 24 times between 2026 and 2030, to 120 quadrillion tokens a month, as companies shift toward AI agents.
| Metric | Trend |
|---|---|
| Individual token cost | Plummeted in recent years (Goldman Sachs, via BBC) |
| Token consumption | Skyrocketed; forecast to rise 24× from 2026–2030 to 120 quadrillion tokens/month |
Enterprises Hit by Unpredictable Bills
Companies and individuals often have only a tenuous grasp on how many tokens they are burning until they run out or receive the monthly bill. BBC reported that Microsoft has reportedly reined back its engineers' use of some third-party coding tools, while Uber tore through its AI coding token budget for a year in a matter of months earlier this year.
Will Venters, Associate Professor of Digital Innovation and Information Systems at the London School of Economics, said companies can be caught out as staff burn through tokens while experimenting with or implementing AI internally.
"People are finding it really hard to manage that cost… it's a non-deterministic output, so it's a non-deterministic value."
Working Around the Meter
Oliver King-Smith, founder of engineering software firm smartR AI, said smaller organizations can "fly under the radar and use [flat fee] personal accounts which I am sure the big vendors don't like."
For enterprise technology buyers, the implication is clear: with Goldman Sachs forecasting 24-fold growth in token consumption by 2030, finance and IT teams need to treat AI usage as a variable cost that can spike without warning. Vendors, meanwhile, are under pressure to recoup hundreds of billions of dollars in LLM investment while unable to predict what their own services will cost to deliver — a tension that is likely to shape pricing models across the AI industry.