NEWS
AI Price Cuts Tighten Power Demand and Rewrite Capacity Deals
OpenAI and Anthropic slash model prices under Chinese pressure; usage jumps while power density and flexible leases become the real constraints for operators.
OpenAI cut the price of its lightweight GPT-5.6 Luna model by 80% on 30 July 2026, to $0.20 per million input tokens and $1.20 per million output tokens, while trimming mid-tier Terra 20%. Anthropic followed days earlier by launching Claude Opus 5 as near-frontier intelligence at roughly half the cost of its top Fable model. Chinese rivals already undercut those prices by several times, and early usage data shows token demand rising faster than prices fell.
The immediate effect lands less on API bills than on the physical layer that runs them. Cheaper tokens pull more inference into GPU-dense halls whose power and cooling already command premiums, while the labs themselves push for flexible capacity deals instead of the decade-long leases hyperscalers once preferred.
Price cuts at the model layer therefore act as a demand shock upstream. Every new agent workflow, coding loop or document pass that becomes viable at the lower rate still needs silicon, power and cooling in a hall that can host it. The race that looks like a software discount is, in practice, a test of who controls dense megawatts on flexible terms.
What OpenAI and Anthropic actually cut
OpenAI’s official update set new API rates for the GPT-5.6 family after efficiency gains in serving. Luna, the fastest and cheapest tier, dropped 80%. Terra, the balanced everyday model, fell 20% to $2 input and $12 output per million tokens. Flagship Sol pricing stayed put, though Fast mode now offers up to 2.5 times standard speed at double the price.
The company framed the moves as passing efficiency improvements to customers so high-volume agent and coding work becomes practical. Customer quotes on the announcement page back the claim: Replit’s Michele Catasta called Luna “the closest we’ve come to intelligence too cheap to meter.” Notion and Ramp reported strong cost-per-task gains.
| Model | Provider | Input / Output per MTok | Change |
|---|---|---|---|
| GPT-5.6 Luna | OpenAI | $0.20 / $1.20 | 80% cut |
| GPT-5.6 Terra | OpenAI | $2 / $12 | 20% cut |
| GPT-5.6 Sol | OpenAI | $5 / $30 (unchanged) | Fast mode added |
| Claude Opus 5 | Anthropic | $5 / $25 | Near Fable at half cost |
| Claude Fable 5 | Anthropic | Higher tier reference | Frontier benchmark |
Anthropic’s Claude Opus 5 at half the frontier price arrived 24 July priced the same as its prior Opus generation yet scored near Fable 5 on coding and knowledge-work suites such as Frontier-Bench and CursorBench. The company made it the default on Claude Max. Both labs kept their absolute frontier offerings expensive while defending the high-volume middle.
The tier design is deliberate. Luna and Opus 5 absorb bulk agent and coding traffic at rates that invite volume. Sol and Fable 5 stay dear so flagship margins and benchmark leadership remain intact. Fast mode on Sol adds a speed premium without touching the base flagship list price, giving heavy users a paid escape valve when latency matters more than unit cost.
- 24 July 2026: Anthropic launches Claude Opus 5 at half the cost of Fable 5 and makes it the default on Claude Max.
- 30 July 2026: OpenAI cuts GPT-5.6 Luna by 80% and Terra by 20%, leaving Sol list pricing unchanged while adding Fast mode.
Taken together, the late-July moves compress the mid-tier bill while leaving the top rung as a scarcity product. That split sets up the volume rebound measured in the weeks that followed.
Chinese models opened a ninefold gap
The pressure came from DeepSeek, Moonshot’s Kimi line and Zhipu’s GLM series. Benchmarking cited across coverage put an equivalent enterprise workload at roughly $544 on Zhipu’s GLM versus $4,811 on Anthropic’s Claude, nearly nine times higher. Artificial Analysis data on DeepSeek’s V4-Flash put average test cost at 3 cents against $3.15 for Claude Fable 5, more than 100 times cheaper in some runs.
- DeepSeek V4-Flash and V4 Pro lead on raw cost per token and cost-per-task indexes.
- Moonshot Kimi K3 ranks high on composite intelligence yet still undercuts Western mid-tiers.
- Zhipu GLM-5.2 sits near the top of open-weight leaderboards at $1.40 / $4.40 per million tokens.
- Open-weight releases let sovereign or third-party hosts run the same models outside US cloud stacks.
Sam Altman noted on X that GPT-5.6 Sol already undercut Anthropic’s Fable 5 on price and token efficiency for many tasks, adding OpenAI would be “happy to deliver at one-quarter of the price” if competition required it. Enterprises already test the cheaper options. Coverage has named DoorDash, Airbnb and coding tools routing some loads to Chinese or open-weight models to control bills.
The ninefold and hundredfold gaps matter less as permanent floors than as bargaining chips. Once a procurement team can show a $544 path beside a $4,811 path for comparable work, Western list prices face continuous pressure even when buyers keep flagship traffic on approved stacks. Open-weight hosts widen that option set further by moving inference off US cloud rails entirely.
| Workload comparison | Lower-cost path | Higher-cost path | Gap cited |
|---|---|---|---|
| Equivalent enterprise workload | $544 on Zhipu GLM | $4,811 on Anthropic Claude | Nearly nine times |
| Average test cost (Artificial Analysis) | 3 cents on DeepSeek V4-Flash | $3.15 on Claude Fable 5 | More than 100 times in some runs |
| Open-weight mid-tier list | Zhipu GLM-5.2 at $1.40 / $4.40 per MTok | Western mid-tiers above that band | Undercuts on list rate |
Usage rose faster than prices fell
TD Cowen analysts examined OpenRouter traffic in the weeks after OpenAI’s cuts. Effective price for Luna fell roughly tenfold while consumption jumped about 14-fold. Terra’s effective price dropped threefold and usage rose fivefold. Revenue from Luna still rose about 34% and from Terra about 45% versus the prior seven days.
Luna usage: ~14× after the cut
Luna revenue: +34% in early window
Terra usage: ~5×
Terra revenue: +45%
That pattern matches the classic efficiency rebound: lower unit cost unlocks tasks that were previously uneconomic, from broader document processing to always-on agents. OpenAI’s own 80% cut on GPT-5.6 Luna pricing was sold exactly on that volume thesis. If the elasticity holds, total tokens processed climb even as price per token drops.
The revenue lift in the early window is the point labs care about most. A tenfold effective-price drop on Luna that still yields a 34% revenue gain means the volume response more than paid for the discount. Terra’s threefold effective-price drop and 45% revenue gain tell the same story at the balanced tier. Elasticity this steep turns a price cut into a utilisation engine, not a margin giveaway.
Power density does not get cheaper
Andrew Batson of JLL has described AI training facilities as needing ten times the power density of traditional data halls and commanding 60% lease-rate premiums. Inference volumes rising on the back of cheaper tokens push more sustained load into the same constrained halls. Memory chips, advanced cooling and grid connections remain short.
JLL’s broader outlook sees roughly 100 GW of new capacity by 2030 and up to $3 trillion in related investment when fit-out is included. The model-layer price war does nothing to ease the physical shortage; it raises the utilisation rate on the scarce power-dense stock that already exists. Operators already treating power as AI’s real brake on capacity will see that constraint tighten if Western volumes rebound.
Some hyperscalers are answering with behind-the-meter generation. Meta’s Alberta project includes a dedicated gas plant for a hyperscale campus, a direct response to grid queues that can stretch years.
Inference is steadier than training spikes, so a 14-fold Luna usage jump does not arrive as a brief peak. It arrives as higher baseload in halls already sold on density premiums. Cooling loops, memory bandwidth and grid interconnects that were sized for yesterday’s token curves become the binding limits long before another API rate card is printed.
Capacity deals move toward flexibility
Frontier labs chasing high valuations must show that falling per-token prices are offset by rising volume. That makes them prefer scalable, shorter-ramp commitments over the fixed 10- to 15-year take-or-pay leases common with hyperscalers. Colocation and neocloud operators negotiating the next 12 to 18 months of deals should expect tougher talks on ramp schedules, exit clauses and utilisation floors even as total megawatts committed grow.
- AI labs and neoclouds push larger footprints with more flexible commencement and termination language.
- Power and cooling availability, not model quality, become the primary underwriting variables.
- Usage-based revenue at the model layer introduces volatility that can affect long-term PPA and build-to-suit credit profiles.
- Operators in the Gulf, Southeast Asia and parts of Europe now price political risk alongside power and latency.
Buyers of colocation should monitor enterprise AI adoption curves as closely as chip shipment data. The infrastructure side, not the public API price list, will decide the next round of contract terms.
Flexible ramps matter because token demand can jump fivefold or fourteenfold in weeks, as the OpenRouter window showed. A lease that only fills on a multi-year schedule leaves money on the table when elasticity spikes, and strands capital when traffic routes elsewhere. Exit language and utilisation floors are how both sides hedge that volatility while still booking larger total megawatts.
Geopolitics still splits the demand map
Export controls already limit which partners can touch the highest Western tiers. Anthropic has moved to restrict certain advanced access for foreign nationals; OpenAI keeps flagship access to approved partners. Dario Amodei has argued that chip restrictions force Chinese labs into less efficient workarounds that consume more power for the same output.
If cost-sensitive workloads keep migrating to Chinese-hosted or open-weight models running on sovereign infrastructure, Western-aligned GPU demand softens. If the price cuts pull enterprise traffic back, US and allied colocation and hyperscale operators capture more of the growth. Chinese firms gaining limited H200 access adds another variable: better local silicon reduces the performance gap that once justified Western premiums.
Data-centre operators caught between regimes already treat political risk as a line item next to land and megawatts. The price war simply makes the fork in demand more visible.
Volume gains still need dense megawatts
The OpenRouter rebound shows why labs accepted deep mid-tier cuts. Luna’s roughly tenfold effective-price drop came with about 14-fold consumption and a 34% revenue rise in the early window. Terra’s threefold effective-price drop came with fivefold usage and a 45% revenue rise. Those ratios vindicate the volume thesis on paper.
They also map straight onto power-dense stock that already carries 60% lease-rate premiums and ten times the density of traditional halls. Tokens that were too expensive to run at scale become steady inference load once Luna sits at $0.20 / $1.20 and Opus 5 holds the near-frontier band at $5 / $25. The bill moves from the API invoice to the hall that must cool and feed the GPUs.
JLL’s path to roughly 100 GW of new capacity by 2030 and up to $3 trillion in related investment assumes that physical build can chase this demand. Behind-the-meter projects such as Meta’s Alberta gas plant show how far operators will go when grid queues stretch years. Until that capacity arrives, utilisation on the existing dense fleet is the relief valve, and it is already tight.
Enterprise traffic splits across cost tiers
DoorDash, Airbnb and coding tools already route some loads to Chinese or open-weight models to control bills. That behaviour will not reverse simply because Luna fell 80% or Opus 5 landed at half of Fable. It will become more granular. Flagship reasoning and regulated workflows stay on approved Western tiers; bulk embedding, draft generation and always-on agents hunt the lowest credible cost-per-task.
Altman’s note that Sol already undercut Fable 5 on price and token efficiency, and that OpenAI would be happy to deliver at one-quarter of the price if competition required it, signals how far list rates can still move. Zhipu’s $544 versus Claude’s $4,811 on an equivalent workload, and DeepSeek V4-Flash at 3 cents against $3.15 for Fable 5, keep a permanent reference point on every procurement spreadsheet.
Open-weight releases add a third rail: sovereign or third-party hosts running the same weights outside US cloud stacks. Export controls and restricted advanced access for foreign nationals shape who may touch the highest Western tiers, but they do not stop cost-sensitive traffic from leaving. Operators underwriting the next 12 to 18 months of deals therefore have to price a demand map that forks by workload class, not only by region.
The next constraint is not the token price
Model prices will keep falling as efficiency and competition continue. The early OpenRouter numbers show volume can more than offset the cuts for the labs. That same volume lands on halls that already run hot and full. Power density, cooling capacity and the flexibility of the capacity contracts that secure them now set the pace more than any API rate card.
Operators who can deliver dense, available megawatts with ramps that match volatile model economics will write the next deals. Those still offering only rigid decade terms will watch the conversations move elsewhere.
-
AI3 months agoOracle Cuts 21,000 Jobs in a Year, Cites AI in 10-K Filing
-
AI2 months agoFable 5 and Mythos 5 Return as US Lifts Anthropic Export Controls
-
AI3 months agoSpaceX’s Google Deal Turns a Rocket Company Into a Cloud Landlord
-
GAMING3 months agoCD Projekt Red Co-CEO: Redemption Arc Isn’t Done, Witcher 4 in 2027
-
CRYPTO3 months agoXPL Rallies 30% Ahead of Plasma One Card Tier Launch
-
NEWS3 months agoGoogle Search Profiles Build a Follow Graph Inside Discover
-
APPS3 months agoDGO App Brings Rs 549 Mobile Pass for FIFA World Cup 2026 in Nepal
-
AI3 months agoMoonshot AI Targets $30 Billion in China’s Fastest AI Funding Sprint
