Connect with us

AI

AI Agents Drive a Fivefold Cost Shock for Enterprises

Gartner says AI agents will lift inference costs more than fivefold through 2028, while most enterprises still sit mid-maturity.

Published

on

Gartner now forecasts inference costs per agentic workflow will rise more than fivefold through 2028, even as token prices keep falling. The warning lands on firms that spent two years treating agents as the cheap way to scale work. Most of them still sit in the middle of the maturity curve, and only about one in five have agentic systems in production.

Gartner analysts Arun Chandrasekaran and Pieter den Hamer framed that bind in June 2026: model races reset every quarter, agents start to run the workflow, and the bill stops looking like a free lunch. The second-order effect is not a smarter chatbot. It is a rewrite of how software is priced and how AI spend is governed.

The Agent Wave Hits the Invoice First

Chandrasekaran’s line in June was blunt. “The AI free lunch is starting to not look so free anymore,” he said. Agents do not send one prompt and wait. They call tools, retry, hand work to other agents, and keep the model busy until a task closes or a human steps in.

That loop is the product pitch and the cost problem. A seat-based SaaS contract assumed a person logged in. An agent does not occupy a seat, yet it can fire large volumes of model requests against the same systems. Finance teams that budgeted a chatbot subscription now face a usage curve they cannot forecast, because they lack meters, caps, and owners.

Token fees are only the visible line. Data plumbing, governance layers, model monitoring, and staff training sit on top of that bill. Firms moving past pilots are finding AI spend harder to predict than a software license, which is the point at which the board starts asking who approved the run rate.

Inference Costs Rise Fivefold Even as Tokens Cheapen

On August 17, 2026, Gartner put a number on the June warning. It said AI inference costs per agentic workflow will increase more than fivefold through 2028. The firm calls this the Inference Paradox: better unit economics that raise the overall cost of AI without a clear path to matching, predictable value.

Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens. There is no reliable, economical one-size-fits-all model on the horizon.

Will Sommer, Sr. Director Analyst, Gartner, August 17, 2026 press release

Sommer’s comparison is the specimen, not a metaphor. A basic chatbot reads a query and answers. An agent has to reason, negotiate, and check itself. Routing a task to an agentic reasoning model, Gartner said, raises provider inference costs by at least five times a chatbot turn, and often much more as the job gets harder. That five-times snapshot is the current gap between a single reply and an agent run. The fivefold figure is the path Gartner draws for the whole workflow through 2028.

THE COST STACK INSIDE AGENTIC WORKFLOWS

  • The 2028 path: Inference cost per agentic workflow rises more than fivefold, even as model token prices keep falling.
  • The live gap: Sending a task to an agentic reasoning model costs at least five times a basic chatbot turn, and more as complexity grows.
  • The cancel risk: Gartner predicted in June 2025 that over 40 percent of agentic AI projects will be canceled by the end of 2027, on cost, unclear value, or weak risk controls.
  • The vendor fog: Of the thousands of firms selling “agentic” tools, Gartner estimates only about 130 are the real thing. The rest is agent washing, chatbots and RPA under a new label.

Sommer’s operational warning is the one that should sit on a CIO slide. Defaulting every job to generic autonomous intelligence, he said, produces unbounded costs orders of magnitude higher than a tuned mix of models. The practitioners arguing on X have been making the same point in plainer language: the expensive agents are often doing work a smaller model, a cache, or a rule could finish. Routing is the control. Token price is the receipt for getting routing wrong.

That is why a market-size chart and a kill-rate can sit side by side without contradiction. Spend can grow while a large share of projects never leave the pilot, because the workflow never got a budget owner or a stop switch.

Only 17 Percent Sit at High AI Maturity

Den Hamer’s maturity split is the reason those bills land badly. Only 17 percent of organizations have reached what Gartner calls a high level of AI maturity, with AI embedded across functions at enterprise scale. A further 51 percent sit at medium maturity, running isolated projects that have not changed how the company works. The remaining 32 percent are still in experiments and proofs of concept.

WHERE ENTERPRISE AI MATURITY STANDS

Maturity band Share of organizations What Gartner says it looks like
High 17 percent AI embedded across business functions, enterprise-scale adoption
Medium 51 percent Isolated projects and departments, no transformational impact yet
Low 32 percent Experimentation and proof-of-concept work

The same shape shows up in agents. Gartner estimates only around one in five organizations have agentic systems running in production. “The remaining 80 per cent is only experimenting with agentic AI or is just thinking about it,” den Hamer said. A separate 2026 Gartner CIO and Technology Executive Survey found more than 60 percent of organizations expect to deploy agents within two years, the steepest adoption curve the firm measured among emerging technologies. Intent is not inventory.

Gartner’s 2026 Hype Cycle for Agentic AI, published April 15, 2026 by analyst Rajesh Kandaswamy, parks the category at the Peak of Inflated Expectations. Most live uses are still narrow, in software engineering, customer support, and operations. Fully autonomous agents, the Hype Cycle says, are not ready for the majority of enterprise jobs. Governance, security, and FinOps for agentic AI now sit on the same curve as the toys, which is an admission that the bill arrived before the operating model.

Why Seat Licenses Still Survive the Agent Pitch

The other unit under stress is the seat. Chandrasekaran and den Hamer were clear that a “SaaSpocalypse” is unlikely. Big SaaS vendors still hold deep workflow integration, domain data, compliance machinery, and customer contracts. What they no longer hold without a fight is the right to bill by headcount while agents do the clicking.

Once an agent completes work without a named user, buyers ask why the price still tracks employee count. That question is already pulling contracts toward consumption, workflow, and outcome fees. Customer service is the early test bed, with some vendors tying price to satisfaction scores and resolved work rather than licenses.

HOW THE SOFTWARE BILL IS BEING REWRITTEN

  • Seat licenses: The old unit, a person with a login, still dominates renewals because it is easy to audit and easy to forecast.
  • Consumption fees: Metered tokens, calls, or agent minutes, which follow the work and make the month-end number jump.
  • Workflow fees: Price attached to a process, a ticket lane, or a job type rather than a named employee.
  • Outcome fees: Price attached to a resolved case, a closed lead, or a satisfaction score, which buyers like and vendors struggle to define.

The shift from seat-based licensing is already a buyer tactic, not a thought experiment. A 2025 BCG IT-buyers survey found 40 percent of buyers name seat reduction as their main lever to cut software spend, and agentic tools make that cut easier to attempt. Vendors then have to defend margin while they reprice the same workflow. Large SaaS firms can absorb that fight. Smaller tools that only ever sold seats will feel it first.

Gartner still sees agents moving into the application layer. It predicts 33 percent of enterprise software applications will include agentic AI by 2028, up from less than 1 percent in 2024, and that at least 15 percent of day-to-day work decisions will be made autonomously by 2028, up from none in 2024. Those are installed-base forecasts, not a death notice for Salesforce-class vendors. They are a notice that the unit of value is in court.

Chinese Open-Weight Models Cut the Token Bill

The cost fight also has a map. Chandrasekaran described a split that has only grown since June. Leading U.S. vendors still push proprietary frontier models. Many of the strongest open-weight systems are coming from China, including Alibaba, DeepSeek, MiniMax, Moonshot, and Zhipu AI. He has seen wide trial and production use of those models across APAC, as firms hunt alternatives to locked-in frontier APIs.

Leaderboards, Chandrasekaran said, “keep changing every quarter,” so the prize for shipping the single most capable model is short. Vendors are crowding into efficiency, reasoning, multimodal work, and agents that can drive software on a desktop. Computer-use and tool wiring are what let a system book, file, and click across enterprise apps. Speech plus reasoning is what puts a voice agent on an airline, bank, or insurer line before a person picks up.

Open weights are the pressure valve on that stack. DeepSeek’s April 2026 V4 preview is the cleanest public marker: the lab put DeepSeek-V4 Preview open-sourced weights on Hugging Face under an MIT license, with a 1-million-token context, a 1.6-trillion-parameter Pro mix (49 billion active), and a 284-billion-parameter Flash mix (13 billion active). The pitch is agentic coding at a cost firms can host or meter themselves. For an APAC buyer staring at a fivefold workflow curve, that is a procurement option, not a research hobby.

It is also a governance option. A model you can run in your own rack is a model whose tokens do not vanish into a vendor invoice you cannot itemize. The trade is operational load, license fine print, and the usual questions about data paths. Firms that already cannot forecast a hosted-agent bill will not magically gain that skill by downloading weights. They will only change which ledger the surprise hits.

Fortune 500 Agent Counts Head Toward 150,000

Cost is what the board sees. Sprawl is what IT inherits. On April 28, 2026, Gartner said an average global Fortune 500 enterprise will have over 150,000 agents in use by 2028, up from fewer than 15 in 2025. Only 13 percent of organizations think they have the right AI agent governance in place.

Max Goss, a Gartner senior director analyst, told the Digital Workplace Summit in London that CIOs are already looking at an ungoverned spread of agents, with risks that include bad information, oversharing, and data loss. Blocking the tools, he said, is not a lasting fix. Staff who cannot work in the approved stack will go around it, and shadow AI is worse than a messy inventory.

That pattern is already visible in hospitals and clinics, where staff running healthcare AI without a map mix official tools with whatever chatbot is closest. Agent sprawl is the same habit with a higher API bill and a machine identity attached.

HOW GARTNER’S AGENT WARNINGS STACKED UP

  1. June 25, 2025: Predicts over 40 percent of agentic AI projects will be canceled by the end of 2027, citing cost, unclear value, and weak risk controls, and estimates only about 130 of thousands of “agentic” vendors are genuine.
  2. April 15, 2026: Places agentic AI at the Peak of Inflated Expectations and reports that more than 60 percent of organizations expect to deploy agents within two years.
  3. April 28, 2026: Projects over 150,000 agents per average Fortune 500 firm by 2028, from fewer than 15 in 2025, with only 13 percent confident in their agent governance.
  4. June 22, 2026: Chandrasekaran and den Hamer warn that the AI free lunch is fading, that only 17 percent of organizations are at high AI maturity, and that only about one in five have agents in production.
  5. August 17, 2026: Forecasts inference costs per agentic workflow will rise more than fivefold through 2028 and names the Inference Paradox.

Goss’s six steps are unglamorous on purpose: write rules for who can build agents, keep a central inventory, give each agent an identity and a retirement date, govern the data it can see, watch what it does, and train people so they stop treating every new bot as a side project. Without that, the 150,000 figure is not a productivity story. It is 150,000 small bills and 150,000 ways to leak a file.

Literacy Separates Pilots From Payoff

The firms den Hamer sees getting value are not the ones with the flashiest demo. They use AI to speed research, harden operations, and lift product quality, and they redesign the process around the model instead of dropping a copilot on last year’s workflow. Auto groups have had to learn the same lesson the hard way, mapping where auto AI actually pays rather than spraying assistants across every desk.

The trait that tracks return, den Hamer said, is literacy. People who know what the tool can do, and what it should not do, get more from it. People who fear it stall the rollout.

Everyone needs to learn about AI. If you do that proactively, we clearly see that ROI then tends to be much higher compared to not (educating employees).

Pieter den Hamer, Gartner analyst, June 2026 comments

Job-security talk has to sit in the same program. Den Hamer argued that psychological safety and a clear story about where AI fits in the work are what let deployments scale. An agent that nobody trusts gets shadowed. An agent that nobody measures gets copied. Both habits feed the sprawl Goss is counting toward 150,000 and the fivefold inference path Sommer put on the 2028 calendar.

Gartner IT Symposium/Xpo runs September 14-16, 2026, on the Gold Coast, with more on AI economics on the agenda. The firms that walk in with a meter, a routing policy, and a training plan are the ones that can use an agent as a worker. The rest will keep paying frontier prices for tasks a smaller model could have finished, then cancel the project when the invoice shows up.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending