Token prices are falling. Most AI budgets are still going up.
That contradiction is at the center of new Gartner research on what the firm calls the inference paradox. Token costs for large language models are projected to fall 95% by 2030. But the total cost of running agentic AI workflows is on track to increase more than fivefold over the next two years.
Why Cheaper Tokens Don't Mean Cheaper AI
The disconnect comes down to volume, not price. As AI systems get more capable, they use dramatically more tokens per task, and the increase is outpacing every efficiency gain on the price side.
Gartner analysts Will Sommer and Sabine Zimmerhansl put it bluntly in their report: "The market is captured by a token-deflation illusion."
Buyers who assume today's falling token prices will show up as savings on next year's AI bill are working from a flawed premise, the analysts argue. Advanced reasoning agents already cost up to 150 times more per task than a basic chatbot.
What's Actually Driving the Cost
A simple chatbot reads a request and returns a reasonable answer. An agent has to reason through the problem, question its own output, adapt when something breaks, and often coordinate with other agents entirely in the background.
That coordination gets expensive fast.
Gartner found that training a medium-sized agentic model with reasoning capabilities costs 2.5 times more than training an equivalent chatbot. Running that model costs 5 times more. And agents typically require 5 to 30 times more tokens than a chatbot to complete the same task.
Multiply that across dozens of agents handling hundreds of tasks an hour, and the numbers stop looking like a rounding error.
Putting a Number on It
Gartner built a Tokenomics Model to price out different types of AI workloads. Basic workflows run around 5 cents per inference token. Summarization and knowledge retrieval cost roughly 10 cents. More complex workflows land near 30 cents, and planning or learning tasks run about 40 cents, roughly 8 to 10 times the cost of a basic workflow.
The upside is real. Gartner has seen AI agents cut customer service response times by 99% in some deployments. The technology genuinely outperforms people at plenty of routine tasks.
The problem is that value and cost aren't rising at the same rate.
How Companies Are Supposed to Respond
Gartner's recommendations mostly come down to discipline. Route simple tasks to cheaper models instead of defaulting to the most powerful one available. Move away from flat compute fees toward usage-based pricing that scales with actual need. Treat every new model release the way a car buyer treats a new car, an asset that starts losing value the moment it's deployed.
Companies should also set hard thresholds before greenlighting any agentic deployment, testing it against token-price swings and compliance costs up front rather than discovering the exposure later.
Gartner's core warning is that businesses defaulting to generic, always-on autonomous AI will end up paying orders of magnitude more than companies that build tightly scoped, optimized systems.
What This Means for Miami
Miami's AI and startup community has leaned into agentic tools quickly, from customer service automation to fintech workflows to real estate and legal research assistants. Most of those budgets were built around the assumption that AI would keep getting cheaper.
Gartner's research suggests the opposite may be true for anything beyond simple chatbot use cases. Founders raising money on the promise of thin AI-driven margins should model their token spend against usage growth, not against today's per-token price, since the two are moving in opposite directions.
For Miami's venture and private equity community evaluating AI-native startups, cost architecture is quietly becoming as important a diligence question as the product itself. A startup that can't explain how it controls inference costs as usage scales may be building a business that gets more expensive to run with every new customer it wins.