Connect with us

AI

OpenAI and Anthropic Slash Mid-Tier Prices as Chinese Models Bite

OpenAI and Anthropic cut mid-tier model prices up to 80% as Chinese open-weight rivals like Kimi K3 and DeepSeek force hybrid stacks and pressure.

Published

on

OpenAI slashed prices on its GPT-5.6 Luna model by 80 percent and Anthropic priced Claude Opus 5 at half the rate of its flagship Fable 5, as Chinese open-weight systems from Moonshot and DeepSeek pull cost-conscious customers away from premium American labs.

The moves arrive after Silicon Data’s token price index fell almost a quarter since mid-July and after companies began capping usage or testing cheaper alternatives. The Financial Times reported the shift on August 14.

The Mid-Tier Got Cheaper Fast

OpenAI’s July 30 update delivered an 80 percent cut on GPT-5.6 Luna, dropping it from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output tokens. Terra, the balanced tier, fell 20 percent to $2/$12. Flagship Sol stayed at $5/$30.

Anthropic launched Opus 5 on July 24 with the pitch of near-Fable intelligence at half the price: $5 per million input tokens and $25 per million output tokens, against Fable 5’s $10/$50. The company also cancelled a planned September price rise on Sonnet 5.

Both labs frame the changes as efficiency gains passed to customers. OpenAI pointed to kernel and serving improvements that cut serving costs. Anthropic stressed that Opus 5 sits in its normal family structure rather than a direct reaction to rivals. A person close to Anthropic told the FT there was “no connection to competitors.”

  • Luna new rate: $0.20 input / $1.20 output per million tokens
  • Opus 5 rate: $5 input / $25 output, half of Fable 5
  • Index move: leading US lab prices down nearly 25% since mid-July per Silicon Data
  • Sol and Fable: premium tiers left largely untouched

Chinese Open Models Made the Math Unworkable

Moonshot’s Kimi K3 and DeepSeek’s V4 Flash and V4 Pro arrived with open weights that developers can download, fine-tune and host themselves. That removes the pure API meter and the need to ship sensitive data to US providers.

Artificial Analysis found Opus 5 at medium effort delivered similar performance and cost per task to Kimi K3 at max effort. GPT-5.6 Luna at max effort matched DeepSeek V4 Flash performance at just under twice the cost per task. On the broader Intelligence Index, Kimi K3 has ranked among the top open models and close to several closed frontier systems.

The open-weight path also benefits from earlier US export and policy pressure that, as covered in reporting on an earlier boost to Chinese open-source models, accelerated domestic Chinese development and global distribution of weights. Developers who once defaulted to closed APIs now route routine volume elsewhere.

On X and in developer forums the conversation has shifted from raw capability scores to tokens per completed task. Sticker prices matter less when a cheaper model burns more tokens or a premium model finishes the job in fewer steps. That efficiency frame is now the ground on which OpenAI and Anthropic must fight.

What the Token Cards Look Like Now

Headline prices still vary by effort setting, caching and how many attempts a model needs. Still, the published rates show the new middle of the market.

Model Input $/MTok Output $/MTok Notes
GPT-5.6 Luna 0.20 1.20 80% cut July 30
GPT-5.6 Terra 2.00 12.00 20% cut
GPT-5.6 Sol 5.00 30.00 Unchanged standard
Claude Opus 5 5.00 25.00 Half of Fable 5
Claude Fable 5 10.00 50.00 Frontier holdout
Kimi K3 (API) ~3.00 ~15.00 Open weights available
DeepSeek V4 Flash ~0.09-0.14 ~0.18-0.28 Very low cost tier

More capable models can finish work with fewer tokens, so a higher sticker price sometimes wins on total bill. Effort knobs let users trade intelligence for cost inside the same model family. That complexity is why procurement teams now run side-by-side evals rather than trusting list prices alone.

Enterprises Already Mix and Match

DoorDash, Airbnb and Siemens are among the firms the FT and follow-on reports named as testing or adopting Chinese models to rein in bills. DoorDash cofounder Andy Fang publicly noted savings by routing lower-level work to a Moonshot model. Some startups have moved entire workloads off Anthropic onto DeepSeek variants.

The pattern is hybrid: frontier closed models for high-stakes reasoning, coding agents or customer-facing work that needs the highest reliability; open or low-cost Chinese models for high-volume classification, drafting, internal search and background agents. Hostinger AI tech lead Mantas Lukauskas called the recent cuts “the first real test” of whether US labs can protect the price of their most advanced offerings.

The US labs have cut the middle and are defending the top.

Lukauskas told the FT that prices for the very best models remain flat to rising. That split is now visible in product defaults too: Opus 5 became the new default on Claude Max and the strongest option on Pro, while OpenAI made Luna far more attractive for background automation and high-volume loops.

Developers also remember moments when American models refused a Chinese alternative filled the gap inside agent systems. Trust, data residency and safety reviews still favor US providers for many regulated workloads, yet the pure cost gap has forced even cautious buyers to dual-source.

The Premium Tier Holds Its Price

Sol and Fable pricing barely moved. Fast or priority modes on the flagship models can even raise effective cost for latency-sensitive work. OpenAI introduced Fast mode for Sol at up to 2.5 times standard speed for twice the price. Anthropic continues to position Fable 5 for the hardest knowledge and coding tasks.

The strategy is clear on paper: keep the margin and the narrative on the models that justify trillion-dollar private valuations, while fighting for volume in the mid-tier where Chinese open weights are strongest. Customer quotes on the OpenAI blog celebrate Luna as “intelligence too cheap to meter” for Replit and as a strong everyday agent for Notion. Those testimonials sell abundance. They do not prove the frontier tier will stay immune forever.

Effort settings and better harnesses let a mid-tier model punch above its list price on many agentic tasks. Once evaluation harnesses and routing layers mature, the volume of work that truly requires Sol or Fable may shrink further. That is the second-order risk the price cuts only delay.

Bills, Caps and the Road to Public Markets

Both labs have pushed more enterprise customers toward usage-based billing and away from flat subscriptions. When bills spike, finance teams impose caps. Uber and others have already rationed tools such as Claude Code. Caps accelerate the search for cheaper substitutes and multi-model routers.

OpenAI and Anthropic are preparing for potential IPOs at valuations that assume sustained pricing power and massive future free cash flow. A lasting price war in the middle of the stack compresses the very margins investors will scrutinize. Chinese open-weight competition also undercuts the idea that only a handful of closed US labs can deliver frontier performance.

Kimi K3 cost-per-task rankings and DeepSeek’s flash pricing show how quickly a new open release can reset expectations. On social platforms the sharper takes note that inference is commoditizing faster than oil or semiconductors ever did. Labs sitting at higher price points are buying time with trust, safety evaluations, enterprise contracts and brand. Time is finite when a $0.20 input model can handle most of a workflow.

The immediate effect of the July and August cuts is relief for customers and a temporary pause in share loss at the mid-tier. The longer test is whether defending the top while discounting the middle can fund the compute build-out and still produce the returns a public market will demand. Hybrid stacks are already the practical answer for many buyers. That reality is now priced into the token index, even if it is not yet fully priced into the private valuations.

US labs still lead on many hard benchmarks and on the safety and compliance packaging large enterprises require. Chinese open models now set the cost floor for everything else. The price war is the market acknowledging both facts at once.

Logan Pierce is a writer and web publisher with over seven years of experience covering consumer technology. He has published work on independent tech blogs and freelance bylines covering Android devices, privacy focused software, and budget gadgets. Logan founded Oton Technology to publish clear, no nonsense tech news and reviews based on real hands on testing. He has personally tested and reviewed dozens of mid range and budget Android phones, written extensively about app privacy, and built and managed multiple WordPress publications over the past decade. Logan holds a bachelor's degree in English and studied digital marketing at a certificate level.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending