Gartner: AI Inference Costs Per Agentic Workflow to Rise More Than Fivefold by 2028
A new forecast from Gartner warns that the economics of autonomous AI are about to get much harder, with per-workflow inference costs surging as agents take on complex multi-step tasks.
The cost of running a single autonomous AI workflow could rise more than fivefold by 2028, according to a forecast from technology research firm Gartner. The projection, published on August 17, 2026, signals a structural shift in the economics of enterprise AI: as systems evolve from answering isolated questions to executing complex, multi-step tasks, the computational bill for each completed job is set to climb sharply. For organizations betting on agentic AI to transform customer service, software development, financial analysis, and back-office automation, the coming years will demand a hard reckoning with infrastructure budgets, token consumption, and return on investment.
The timing is significant. The industry has spent the past three years racing to deploy AI agents across enterprise applications. Gartner previously estimated that 40% of enterprise applications would feature task-specific AI agents by 2026, up from less than 5% in 2025. That adoption curve is now colliding with a harsh reality: the more capable and autonomous an agent becomes, the more inference it requires. The fivefold cost increase is not a worst-case scenario but a central forecast, and it arrives at a moment when many organizations are already questioning whether their AI investments are paying off.
The Mechanics of a Fivefold Increase
The driver behind the projected cost surge is not a single technology but a structural change in how AI systems operate. A traditional AI assistant handles one prompt and returns one response. An agentic workflow, by contrast, may involve dozens or even hundreds of sequential inference calls. A single task—such as researching a market, drafting a report, and executing a series of follow-up actions—can trigger a chain of model queries, each consuming thousands or millions of tokens. Gartner's analysis points to this multiplication of inference demands as the primary engine of cost growth.
Compounding the issue is the rising complexity of the models themselves. Larger context windows, more sophisticated reasoning capabilities, and multi-modal inputs all increase the computational load per inference. When these capabilities are strung together in an autonomous workflow, the cumulative token consumption grows far faster than the number of steps would suggest. The result is that a workflow that costs a few cents today could cost several times more by 2028, even if the underlying price per token declines. Efficiency gains in hardware and model serving may partially offset the trend, but Gartner's forecast implies they will not come close to neutralizing it.
The Broader Spending Landscape
The fivefold increase in per-workflow inference costs sits within a larger surge in AI infrastructure investment. Gartner separately forecast that worldwide spending on AI-optimized infrastructure-as-a-service will reach $42 billion in 2026, a 96% increase from the previous year. That near-doubling of IaaS spending reflects not only more AI workloads but also the migration of those workloads to specialized hardware and cloud environments designed for high-throughput inference. Organizations are effectively paying a premium to run agentic systems at scale, and that premium is expected to grow.
The economic pressure extends beyond infrastructure. Gartner has also predicted that AI coding costs will surpass the average developer's salary by 2028, driven by surging token consumption in software development agents. This is a striking data point for engineering leaders: a tool intended to augment or replace human developers could, in some scenarios, cost more than hiring the developer it was meant to displace. The same logic applies across other domains. A customer service agent that handles thousands of interactions per day may generate inference bills that rival the cost of a human support team, especially when the agent is expected to perform multi-step resolutions rather than simple FAQ lookups.
Canceled Projects and the ROI Reckoning
The financial strain is already producing casualties. Gartner has projected that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs and unclear business value as the primary reasons. This is not a prediction about technical failure; it is a prediction about economic failure. Many organizations launched agentic AI initiatives during a period of experimentation and optimism, often without a rigorous model for measuring the total cost of ownership. As inference bills climb and the fivefold increase materializes, finance teams are likely to intervene.
The cancellation forecast should be read as a warning to executives and founders. The barrier to AI adoption is shifting from technical capability to economic feasibility. Two years ago, the central question was whether an AI agent could perform a given task. Today, the more important question is whether it can perform that task at a cost that justifies the outcome. For many use cases, the answer may be no—at least with current model pricing and workflow designs. Organizations that fail to account for the fivefold increase in their long-term budgets risk not only wasted investment but also operational disruption when projects are abruptly halted.
What Leaders Should Do Differently
The forecast does not mean agentic AI is doomed. It means the approach to deploying it must change. Gartner's analysis implies that cost optimization will become a core competency for AI teams, alongside model selection and prompt engineering. That could mean choosing smaller, task-specific models over large general-purpose ones for high-volume steps in a workflow. It could mean caching intermediate results, batching inference calls, or designing agents to make fewer, more deliberate model queries. It could also mean rethinking which workflows are truly worth automating with autonomy, as opposed to simpler retrieval or rule-based systems.
For CTOs and technical founders, the fivefold figure provides a concrete planning benchmark. Budgets for agentic AI should not assume linear scaling with usage; they should assume a steep, compounding curve. Procurement teams negotiating with cloud providers will need to scrutinize inference pricing more carefully than ever. And boards evaluating AI strategy should demand the same financial rigor for agentic workflows that they would apply to any major capital expenditure. The era of treating AI inference as a negligible marginal cost is ending.
Looking ahead, the next two years will separate organizations that treat AI economics as a first-class design constraint from those that treat it as an afterthought. The fivefold increase in per-workflow inference costs is not a distant possibility—it is a forecast rooted in current adoption trajectories and architectural trends. Companies that proactively redesign their agentic systems for cost efficiency, measure ROI with precision, and set realistic expectations about what autonomous AI can deliver at an acceptable price will be positioned to benefit. Those that ignore the coming cost curve may find themselves among the 40% canceling projects by the end of 2027, having learned an expensive lesson about the difference between what AI can do and what it is worth.
Sources
Written by an AI editorial process from the sources above. Errors may occur.
Newsletter
Get the AI news that matters
One short brief with the day's most important AI stories — written for professionals.
We send a confirmation link. No spam. Unsubscribe anytime.
Read next
AI Video Generation Moves From Research Toy to Production Infrastructure
Diffusion models now generate coherent 4K footage at a marginal cost of €0.02 per second, reshaping advertising, film previsualization, and stock footage economics.
29 Aug 2026
Nvidia Shatters Expectations Again, but Investors Still Waver
Record $96.2 billion quarterly revenue and a $500 billion AI infrastructure financing push underscore Nvidia's dominance, even as its stock slips on sky-high expectations.
29 Aug 2026
DeepSeek's V4-Pro-0813 Caps a Three-Year Sprint to Frontier AI
The Hangzhou lab's latest model, with a 1M-token context and 384K-token output, signals a structural shift in who builds and distributes the world's most capable AI systems.
29 Aug 2026