AI
Agentic Workflows Turn Cheap Tokens Into Larger AI Bills
Token prices collapsed at the commodity end, but agentic workflows still climb to frontier models, so AI bills rise even as unit costs fall.
Token prices for GPT-3-class quality fell about a 1,000 times in three years, yet agentic AI invoices keep climbing. Compiled figures from a16z, Epoch AI, and Stanford’s AI Index put that class near $60 per million tokens in late 2021 and about $0.06 by late 2024. The slide in the budget deck is not lying about the collapse. It is measuring the wrong unit.
What you buy is a workflow, and the workflow is built to spend this year’s models, retries, and tool loops faster than last year’s rates fall. That gap is the inference paradox: the unit price drops, the invoice rises, and both numbers can be true.
Last Year’s Intelligence Is Almost Free
Stanford’s 2025 AI Index tracked a narrower slice and still found a cliff. GPT-3.5-level quality on MMLU went from $20 per million tokens in November 2022 to $0.07 by October 2024, more than 280-fold in about a year and a half, with Gemini-1.5-Flash-8B at the low end. Epoch AI’s wider read is that inference costs have been falling anywhere from nine to 900 times a year, depending on the task. Goldman Sachs Research, in a May 2026 note from analyst Jim Schneider, said semiconductor suppliers are still cutting cost per token by 60 to 70 percent a year.
The cheap end of the board is real. The frontier is not. Anthropic’s price list still carries Claude Opus 5 at published rates of $5 and $25 per million input and output tokens, the same sticker it used for Opus 4.8, with Claude Fable 5 and limited-availability Mythos 5 at $10 and $50. Those are the models an agent climbs to when the small one fails, and they sit in the same neighborhood as GPT-3 launch pricing even after the commodity collapse.
Opus 5 shipped on July 24, 2026, with thinking on by default, a 1 million token context window, and a Fast mode that runs about 2.5 times quicker at $10 and $50. Cache hits on most Claude models price at one-tenth of base input. Fable 5.1 cache hits are $0.25, a quarter of Fable 5’s $1 cache-hit rate. Batch jobs take 50 percent off. The discounts are real. They still sit under a frontier list that never fell with last year’s models.
Introducing Claude Opus 5.
It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price. pic.twitter.com/GQWhcq2CQL
— Claude (@claudeai) July 24, 2026
Teams that already hit quota burn on an earlier Opus model know the pattern: the same dollar rate, more work per call, and a default that thinks until you stop it.
THE ANTHROPIC RATE CARD
| Model | Input / million | Output / million | Cache-hit input |
|---|---|---|---|
| Claude Fable 5 | $10 | $50 | $1 |
| Claude Opus 5 | $5 | $25 | $0.50 |
| Claude Sonnet 5 | $2 | $10 | $0.20 |
| Claude Haiku 4.5 | $1 | $5 | $0.10 |
Sticker prices per million tokens are a weak comparison across labs once tokenizers differ, so the useful test is still cost per finished task on your own traffic. The table is what the agent sees when it escalates.
The Retry Loop That Multiplies a Cheap Call
Heming Fu, Shan Lin, Qianqian Xie, and Guojun Xiong gave the gap a name in a July 2026 paper: token inflation, the ratio of true workflow cost to a single-call estimate. On multi-hop question answering, a Qwen2.5-7B model reached token inflation as high as 4.25 times, with only 21.4 percent accuracy. GPT-4o still inflated 2.92 times on the same HotpotQA split, and 1.31 times on GSM8K. FrugalGPT, a standard cost-aware router, underestimates true expense by more than 2 times on hard tasks; on 200 hard GSM8K queries the authors put that miss at 110 percent.
The mechanism is ugly in a simple way. A failed reasoning chain is not only a wasted call. It is a call that gets sent again, history attached, often to a more expensive model. Fu’s group found that forwarding a failed Qwen2.5-7B chain to GPT-4o cut accuracy from 73.9 percent to 39.1 percent on the failure set, a drop of 34.8 percentage points. Their InflationAgent router, which discards the failed chain and picks models by expected accuracy over predicted true cost, hit 94.7 percent on GSM8K against FrugalGPT’s 91.0 percent, and needed 31 percent fewer tokens to match FrugalGPT’s accuracy.
A nominally cheap model that retries four times is not cheap. It is a full-price ticket with a worse answer at the end.
Why Tool Calls Leave GPUs Sitting Idle
Ji-in Kim, Byeongjun Shin, Jinha Chung, and Minsoo Rhu at KAIST measured the same loops against a serving stack rather than a budget line. Tool-augmented agents issued 9.2 times as many model calls as a chain-of-thought baseline. A tree-search agent, LATS, averaged 71.0 calls per single request. GPU memory per request rose 3.0 times on average and 5.4 times in the worst case, because each turn appends reasoning and tool output to the next prompt.
WHAT AGENTS DO TO A SERVING STACK
- Call volume: Tool agents run 9.2 times as many model calls as chain-of-thought; LATS averages 71.0 calls per request.
- Idle time: On HotpotQA and MATH, GPU idle periods waiting on CPU or external tools reached 54.5 percent of execution.
- Memory: KV-cache memory per request ran 3.0 times a CoT baseline, and 5.4 times at the high end.
- Energy: GPU energy per query rose 62.1 times to 136.5 times versus single-turn ShareGPT inference.
Prefix caching cut end-to-end agent latency 15.7 percent and trimmed LATS memory 64.8 percent by reusing shared prefixes, which is one of the few serving discounts that shows up before the invoice. The rest of the stack still pays for GPUs that sit dark while a search tool or a code runner comes back. That idle time does not appear on a per-token rate card, and it is why a CX slide that only plots input price misses the bill the data-center team actually defends.
Extra Tokens Start Flipping Correct Answers
There is a second leak, and it is worse than waste. Shu Zhou, Rui Ling, Junan Chen, Xin Wang, Tao Fan, and Hao Wang tested what happens when you keep spending on reasoning in When More Thinking Hurts. Marginal utility on their set dropped from +1.8 percent per 500 tokens in the 2,000 to 4,000 range to +0.1 percent between 8,000 and 12,000, then turned negative. Past about 7,000 tokens the model flipped more previously correct answers to wrong than wrong answers to right. Their 32B run peaked at 55.8 percent accuracy at 12,000 tokens and fell to 54.9 percent at 16,000.
FLIP RATIOS AS THE BUDGET GROWS
| Reasoning budget | Flip ratio (wrong over right) |
|---|---|
| 7,000 tokens | 1.09 |
| 8,000 tokens | 1.42 |
| 12,000 tokens | 3.29 |
| 16,000 tokens | 7.55 |
At 16,000 tokens the model was 7.55 times more likely to wreck a good answer than to fix a bad one. On easier MATH-500 Level 1 and 2 items, the flip ratio crossed 1.0 near 2,000 tokens. Easy tickets are most of a service queue, and they hit the wall first. The Efficient Agents work comes at the same point from the other side: a leaner design kept 96.7 percent of the accuracy at $0.228 per problem against $0.398 for richer systems. More compute is a default, not a plan.
Salesforce Already Repriced the Conversation
None of this would bite as hard if the contract named the work. It usually does not. Live CX buying still splits across six units, and each one hides the multiplier in a different place.
SIX LIVE CX PRICING UNITS
- Per seat: Tokens are bundled into a user license, so the invoice looks stable until usage explodes inside the seat.
- Per channel: Voice and chat carry different meters; Amazon Connect lists $0.038 a voice minute and $0.010 a chat message.
- Per component: Each bot, flow, or add-on bills on its own clock.
- Credits: Salesforce sells $500 per 100,000 Flex Credits, about $0.10 a standard action at 20 credits, after the $2 conversation unit stopped fitting the work.
- Per action: You have to model the whole workflow before you can forecast a quarter.
- Per resolution: You pay when the bot closes the ticket; Zendesk describes the idea and does not publish a rate.
Genesys still runs about $75 to $240 per user per month. Agentforce Service is listed around $125 per user per month on top of a Service Cloud seat. Help Agent resolutions draw 400 Flex Credits, which is $2 at the list credit price, a flat outcome charge that does not move with retries. Per-message looks cheap until the agent creates a second contact and you pay twice for failing once. Per-resolution looks aligned until you notice it pays the vendor not to escalate. Per-action asks a CX lead to become a workflow engineer before the quarter starts.
Salesforce launched Agentforce at $2 a conversation, then added Flex Credits in May 2025 after that unit stopped matching the job. A three-action ticket that costs $0.30 on credits still costs $2 on the old conversation meter. The reprice is the tell: the industry already admitted the unit of work was wrong, then replaced it with another unit that still is not a token, a retry, or a failed chain.
Goldman Projects 120 Quadrillion Tokens a Month
Schneider’s May 2026 note is not an argument against spending. Cheaper units drive more use; that is Jevons, and it is what happened to coal, steel, and bandwidth. Goldman expects token consumption to multiply 24 times between 2026 and 2030, to 120 quadrillion tokens a month, and 55 times by 2040 if enterprise agents hit peak adoption. Daily AI queries go from about 5 billion to about 23 billion by 2030 in the same note. Query count grows slower than token count, which is the agent loop showing up in the macro number.
Gartner put the product-level version on paper on August 17, 2026. Will Sommer, a senior director analyst, said inference costs per agentic workflow will increase more than fivefold through 2028. A March 2026 Gartner analysis had already put agentic tasks at 5 to 30 times more tokens than a standard chatbot, and said routing a task to an agentic reasoning model raises provider inference cost by at least five times, often more as the task gets harder. The same research line still forecasts that inference on a one-trillion-parameter model will cost providers over 90 percent less in 2030 than in 2025. Falling unit cost and a rising workflow bill are the same forecast.
Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens.
Will Sommer, Senior Director Analyst, Gartner newsroom, August 17, 2026
Spending more on inference in 2029 than you do in 2026 is not, by itself, mismanagement. The failure is narrower. It is a bill you cannot attach to a workflow, an agent, or a closed ticket.
Put the Meter in the Contract First
The controls that pay back fastest are boring, and they work before you scale.
FIVE CONTROLS THAT PAY BACK FIRST
- Instrument the pilot: Per-agent and per-workflow token telemetry with alerts, from day one. One healthcare deployment ran from $12,000 to $68,000 a month over six weeks on a retrieval fault that sat unnoticed for two of them.
- Cap the loops: Hard retry ceilings, mandatory human handoff at the limit, and a fresh-escalation rule that drops the failed chain. Fu’s 34.8 point drop is what you buy when you forward the wreckage.
- Budget the thinking: Per-task reasoning caps, set by difficulty, not a uniform ceiling. Zhai and colleagues got up to 12.8 percent better accuracy on MATH at the same budget by varying compute per item. Above Zhou’s crossover you are paying list price for a worse answer.
- Take the engineering discounts: Cache reads at a tenth of standard input, compact context, retrieve just in time, and keep tool schemas off the cheap calls. One team took $40,000 a month to $24,000 on routing alone. Some teams also move steady volume onto purpose-built cheaper inference chips instead of sending every step to a frontier API.
- Name the unit in writing: Cached tokens, tool execution, failed calls, and retries, plus audit rights against your own telemetry. Current clause guidance is 120 days’ notice before a reprice or a redefinition of the meter; on a mid-year meter, six months is the ask. Then score value per thousand tokens against agreed outcomes, not against volume.
Organizations that do this report spending 60 to 70 percent less for equivalent output. The tokens will keep getting cheaper. The invoice is still the part you can name, cap, and put in the contract.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
