Dive Brief:
- As enterprise AI use evolves from assistive features to multistep processes, tech leaders will face growing costs for AI inference, a report from Gartner released Monday found. AI inference costs are projected to increase more than fivefold through 2028, according to the analyst firm.
- Tokens are becoming more cost efficient, but changing AI pricing models and more complex workflows are posing challenges for tech leaders in charge of cost management. The rate of innovation is currently outpacing the cost savings enterprises experience from their AI providers, Gartner’s data showed.
- Product leaders won’t be able to rely on more efficient tokens to keep costs in check for their organizations, Will Sommer, senior director analyst at Gartner, said in the report. “Each successive generation of AI capability will necessitate more, and often more expensive, tokens,” he said. “There is no reliable, economical one-size-fits-all model on the horizon.”
Dive Insight:
AI tokens may be getting more cost effective for enterprises, but organizations’ reliance on AI systems and how much they use the technology will keep driving up spending.
Enterprises have worked to deploy agentic systems in 2026, with spending on AI-optimized infrastructure as a service — the compute that supports large language model training and operation — projected to nearly double through the end of 2026, reaching $42 billion, Gartner found earlier this month.
Sommer told CIO Dive in a July interview that Gartner estimates global token consumption to be about 300 trillion tokens per day, an annualized growth rate of 500% to 600%.
Hard-to-predict AI costs are forcing some enterprises to rethink their implementation plans. Nearly half of organizations reported that they’ve escalated AI spending surprises to the board, according to a July report by cost management software Mavvrik.
Three main factors are shaping current token economics, Gartner’s Monday report found. Foundational model costs are going down — a win for enterprises — but improved AI efficiency is unlocking more powerful, expensive models to chase higher-value AI applications. Those sophisticated AI workflows use far more tokens than earlier models’ simple chatbot interactions, which is driving up overall inference costs.
The rate of innovation in AI is outpacing the cost curve, Gartner found. The analyst firm calls it the "Inference Paradox," a system where the overall cost of AI escalates without providing a clear pathway to predictable value.
“The harsh economics of the Inference Paradox are exemplified by the differences between a simple chatbot and an AI agent,” Sommer said in the report. “Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate and question itself.”
Achieving ROI from more advanced systems like agentic AI models or reasoning agents requires much higher returns than basic models, Gartner’s report said. Otherwise, companies will need to highly optimize their models to complete complex tasks relative to cost-effective intelligence, Sommer said.
Both options are possible, but require significant overhauling of business workflows, Sommer said in July.
“Every CIO is going to have different priorities, they're going to have different things that are mission-critical to the organization,” Sommer said. “The CIO has to work with the rest of the organization to determine what the priorities are, and whether or not AI is a cost efficient way to solve those problems.”