AI
Cheaper AI Tokens Shift Risk Onto GPU Campus Landlords
OpenAI’s 80% Luna cut and Anthropic’s half-price Opus 5 chase Chinese labs, while 10x denser halls and 20-year power leases set who carries the risk.
OpenAI cut GPT-5.6 Luna by 80% on July 30, to $0.20 per million input tokens, after Chinese labs priced many jobs at a fraction of Western rates. Anthropic answered six days earlier by shipping Claude Opus 5 at half of Fable 5. The cheaper stickers are a volume bet. The constraint that still prices the next hall is power, cooling, and who carries a 20-year occupancy risk.
That is the AI price war as it hits data centre demand. Token rates can fall in a week. A GPU campus cannot.
OpenAI Cuts Luna Prices by 80%
On July 30, OpenAI said it was passing efficiency gains through to customers, with 80% less for GPT-5.6 Luna and 20% less for GPT-5.6 Terra. Luna, the fast everyday model, moved to $0.20 per million input tokens and $1.20 per million output tokens, from $1 and $6 at launch. Terra, the mid-tier workhorse, moved to $2 and $12, from $2.50 and $15. Sol, the flagship, stayed at $5 and $30 on that day, while Fast mode replaced Priority Processing at twice the standard price and up to 2.5 times the speed.
The company said GPT-5.6 Sol had rewritten production kernels inside a human-led loop, cutting the end-to-end cost of serving the model by 20%, and that its experiments lifted token-generation efficiency by more than 15%. Those gains, OpenAI argued, make high-volume work cheaper to run on the same silicon. Luna, it said, now delivers performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task, and at nearly nine times the speed. On professional work measured by Agents’ Last Exam, OpenAI put Luna’s estimated cost per task nearly 99% below Claude Fable 5.
Michele Catasta, president and head of AI at Replit, said Luna was “the closest we’ve come to intelligence too cheap to meter.” Sid Pardeshi, CTO and co-founder of Blitzy, said Luna handled 2.2 times more context with 8.5 times fewer output tokens, at 87% lower cost than GPT-5.4 mini, across thousands of production calls. The pitch is abundance: more jobs on the same cluster, if quality holds.
THE PRICE CUTS SINCE JUNE
- June 9, 2026: Anthropic ships Claude Fable 5 at $10 per million input tokens and $50 per million output tokens.
- June 26, 2026: OpenAI previews GPT-5.6 Sol, Terra, and Luna, with Sol listed at $5 and $30, about half of Fable 5 on input.
- July 14, 2026: Sam Altman says Sol is already cheaper and more token-efficient than Fable on many tasks, and that OpenAI would go to one-quarter of the price.
- July 16, 2026: TeraWulf announces a 20-year, $19 billion Anthropic lease for 401 MW in Kentucky.
- July 24, 2026: Anthropic releases Claude Opus 5 at $5 and $25, half of Fable 5.
- July 30, 2026: OpenAI cuts Luna 80% and Terra 20%, and adds Sol Fast mode.
Subscription prices for ChatGPT and Codex did not move. Terra and Luna now burn fewer credits against the same quotas, which is another way of saying the same seat can generate more tokens before it hits a cap.
Opus 5 Is Anthropic’s Half-Price Volume Bet
Anthropic did not cut the Opus sticker. It raised what that sticker buys. Claude Opus 5 landed on July 24 at $5 per million input tokens and $25 per million output tokens, the same as Opus 4.8, and exactly half the price of Fable 5. The company called it a thoughtful, everyday model that comes close to Fable 5’s frontier intelligence, and made it the default on Claude Max and the strongest model on Claude Pro.
It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.
Anthropic, Introducing Claude Opus 5, July 24, 2026
On Frontier-Bench v0.1, Anthropic said Opus 5 more than doubles Opus 4.8 at a lower cost per task. On CursorBench 3.2 at max effort, it lands within 0.5% of Fable 5’s peak score at half the cost per task. On ARC-AGI 3, Anthropic reported a score three times the next-best model. Scott Wu, chief executive of Cognition, said that on FrontierCode 1.1, “Claude Opus 5 approaches Fable-level performance at half the cost.”
That is a volume argument dressed as a model card. If near-frontier work can run at Opus rates, more of it will. Fast mode for Opus 5 doubles the price again, to $10 and $50, for 2.5 times the default speed, which is the same trade OpenAI is selling on Sol: pay more only when latency is the product.
Sam Altman had already framed the contest as a race to the bottom of the useful price. On July 14 he wrote that GPT-5.6 Sol was half the price of Fable and about twice as token-efficient on many jobs, and that OpenAI would be happy to deliver at one-quarter of the price if the market demanded it.
GPT-5.6 sol is half the price and ~twice as token efficient as fable in many cases for accomplishing the same task.
happy to deliver at one-quarter of the price.
— Sam Altman (@sama) July 14, 2026
Tokens Per Job Break the Rate Card
Chinese labs are why those cuts arrived as a cluster. DeepSeek, Moonshot’s Kimi, and Zhipu’s GLM have closed enough of the quality gap on common enterprise work that a Western premium is harder to defend. One benchmarking exercise put an equivalent job at about $544 on Zhipu’s GLM and $4,811 on Anthropic’s Claude, a gap of nearly nine times. List prices tell the same story in plainer units. DeepSeek’s V4 Flash sits near $0.14 input and $0.28 output per million tokens. GLM-5.3 lists at $1.40 and $4.40. Kimi K3, the dearer open-weight flagship, is $3 and $15, still under Fable 5 and in the same band as OpenAI’s mid-tier Terra on output.
LIST PRICES PER MILLION TOKENS
| Model | Input | Output | What moved |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | 80% cut on July 30 |
| GPT-5.6 Terra | $2.00 | $12.00 | 20% cut on July 30 |
| GPT-5.6 Sol | $5.00 | $30.00 | Unchanged on July 30 |
| Claude Opus 5 | $5.00 | $25.00 | Half of Fable 5 |
| Claude Fable 5 | $10.00 | $50.00 | Anthropic flagship list |
| DeepSeek V4 Flash | $0.14 | $0.28 | Chinese workhorse list |
| GLM-5.3 | $1.40 | $4.40 | Zhipu flagship list |
| Kimi K3 | $3.00 | $15.00 | Moonshot flagship list |
OpenAI publishes the complete API pricing details for the GPT-5.6 family, including cached-input and batch lanes that pull the effective rate lower still. The sticker is still not the invoice. A model that lists a cheaper output rate can burn more tokens per finished job, then retry, and the bill flips. Teams that actually run review and coding agents already treat Opus 5 as a precision lane inside a mix, not as the default on every ticket, because extra context and extra chatter swamp a 20% gap on the rate card. Open-weight models such as GLM-5.3 and Kimi K3 now handle a large share of bulk work, which is why there is no single default winner on price or quality. The Western cuts are an attempt to pull that bulk back onto US-hosted GPUs.
JLL’s 60% Premium on AI Training Halls
None of that cools a rack. Andrew Batson, global head of data centre research at JLL, has been blunt about the physical gap between a classic data hall and an AI training site.
We’re witnessing the emergence of an entirely new infrastructure paradigm where AI training facilities demand 10x the power density and command 60% lease rate premiums over traditional data centers.
Andrew Batson, Global Head of Data Centre Research, JLL
JLL’s early-2026 outlook put installed capacity on a path from 103 GW in 2024 to 200 GW by 2030, with AI workloads rising from about 25% of global capacity in 2025 to 50% by 2030. It described a five-year investment surge of up to $3 trillion, including $1.2 trillion of new real estate value and about $870 billion of new debt. Global occupancy was 97%, and 77% of the construction pipeline was already committed. The firm expects a turn in 2027, when inference overtakes training as the main AI use case, which is exactly the workload cheaper tokens are built to multiply.
JLL’S PHYSICAL CONSTRAINTS
- Power density: AI training halls need about 10x the power of a classic data hall and take 60% lease premiums.
- Pipeline: Occupancy sits at 97%, with 77% of new build already spoken for.
- Grid wait: Average connection times in primary markets now exceed four years.
- Kit lead time: Equipment averages 33 weeks, about 50% above pre-2020, and more than half of 2025 projects slipped by at least three months.
Matt Landek, JLL’s global division president for data centres, said hyperscalers were allocating $1 trillion of data centre spend between 2024 and 2026, against supply limits and those four-year grid delays. Cheaper Luna jobs and half-price Opus 5 calls do not create substations. They fill the halls that already have power, and they make the next vacant megawatt more valuable.
Anthropic Locked 401 MW in Kentucky for 20 Years
Watch the leases, not the blog posts. On July 16, TeraWulf said it had signed Anthropic to a 20-year agreement for about 401 MW of critical IT load at the Justified Data campus in Hawesville, Kentucky, a former Century Aluminum smelter. Contracted revenue is about $19 billion over the first term. Two five-year renewal options can stretch occupancy to 30 years. Initial capacity is due in the second half of 2027, with the full 401 MW ramped by early 2028. Rent starts as each slice of the premises is delivered, which is a ramp schedule hiding inside a long shell.
The site had drawn about 480 MW when it was a smelter, with multiple high-voltage lines and an onsite substation already live. TeraWulf paid $200 million in cash and reported total acquisition consideration of about $301.9 million after a 6.8% minority interest and costs. In March 2026 it put a $500 million delayed-draw construction facility behind the campus. Paul Prager, TeraWulf’s chairman and chief executive, said the lease “establishes a long-duration revenue stream with one of the world’s leading AI companies.” Anthropic’s payment obligations are expected to receive investment-grade credit support, which is the landlord’s answer to a tenant whose product now bills by the token.
That campus is one slice of a much larger lock-up. Announced Anthropic compute leases since October 2025 cover about 14.8 GW. The company has also described plans for $50 billion in American computing infrastructure, including purpose-built sites with partners. Those are not 90-day cloud reservations. They are decade-scale claims on power. The cheap API and the long lease are the same bet, seen from opposite sides of the meter: fall in price per token, rise in tokens, and a building that has to be paid for whether the next model is Luna or Fable.
Who Carries Occupancy Risk After the Cuts?
Usage-based billing makes lab revenue jump around with token mix, effort settings, and which model a customer actually routes. A 20-year rent check does not jump around. Landlords and lenders now underwrite that mismatch. Credit wraps, parent guarantees, and phased delivery are how a token business borrows a utility’s time horizon. Buyers of colocation should expect AI-lab tenants to push harder on ramp dates and exit language even as they ask for bigger footprints, because the volume story only works if unused halls can be slowed and full halls can be filled fast.
Geopolitics sits on the same term sheet. Export controls already shape who may run the top Western models. Anthropic had a narrow window to pull Fable 5 access for foreign nationals, and OpenAI’s GPT-5.6 flagship tier opened first to a limited set of approved partners. Dario Amodei, Anthropic’s chief executive, has argued that keeping advanced chips from Chinese labs would force those labs to spend more power and more engineering to match Western output. Operators in the Gulf, Southeast Asia, and parts of Europe now price political risk next to power and fibre, because a customer that can move inference to a sovereign or open-weight stack is a customer that can strand a hall.
If the Western cuts pull enterprise work back onto US-aligned campuses, GPU-dense colocation stays bid. If cost-sensitive jobs keep moving to Chinese-hosted or open-weight models running on local iron, the demand curve for those Western halls bends, and the long lease becomes the problem instead of the prize. Shops already split the difference. Hard cases stay on Sol, Fable, or Opus 5. Bulk drafting, classification, and background agents slide to Luna, GLM, DeepSeek, and Kimi. That split is what a capacity contract now has to survive.
Grid Queues Still Run Past Four Years
JLL’s global lease-rate forecast is a 5% compound annual rise through 2030, and 7% in the Americas, where grid bottlenecks are tightest. Average equipment lead times of 33 weeks, four-year interconnection queues, and a 2025 slip rate above 50% mean the next 12 to 18 months of deal-making will be about delivery dates more than about which model is cleverer. Power and cooling, not benchmark rank, cap growth.
WHAT THE NEXT CAPACITY TALKS WILL HINGE ON
- Ramp, not just term: Rent that starts as each hall is delivered, with slower fill rights if token demand lags.
- Credit wraps: Investment-grade support on lab payment obligations, because usage billing is jumpy and a 20-year shell is not.
- Exit and contraction: Harder bargaining on unused megawatts, even inside a large committed footprint.
- Power as the scarce clause: Substations, liquid cooling, and grid dates outrank model-quality reps in the term sheet.
- Two products at once: Long campus leases for locked power, plus shorter merchant GPU for overflow, priced as separate risks.
Frontier labs chasing high valuations have to show that falling prices per token are offset by rising volume. That is why they still sign 20-year Kentucky paper and still want flexibility on the way in. The $19 billion Hawesville lease does not kill the volume story. It is the volume story, poured into concrete, with rent that only starts when each block of 401 MW comes alive. Token stickers will keep moving. The queue for the next transformer will not.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
