AI
OpenAI’s 80% Luna Cut Multiplied Token Use Tenfold
OpenAI’s 80% Luna token cut produced roughly tenfold more use, so cheaper ChatGPT API rates did not shrink enterprise AI bills.
OpenAI cut GPT-5.6 Luna token prices 80% on July 30, and its finance chief later said use rose roughly tenfold. The June talk of cheaper ChatGPT prices was not a rumor that died in the filing drawer.
The discount landed on the cheap, fast model first. Enterprise revenue at OpenAI still climbed, because traffic filled the gap the list price left.
The 80% Luna Cut That Followed the June Leak
People familiar with the matter said in June that OpenAI was weighing cuts to the prices it charges for tokens, the billing units for prompts and replies, days after a confidential S-1 went to the U.S. Securities and Exchange Commission on June 8. Chief executive Sam Altman had already called AI bills “a huge issue” for customers and said the company would find ways to deliver “more value for less spend.”
The first hard move came seven weeks later. On July 30, three weeks after GPT-5.6 went to general availability, OpenAI said GPT-5.6 Luna will cost 80% less, taking the API rate to $0.20 per million input tokens and $1.20 per million output tokens, down from $1.00 and $6.00. GPT-5.6 Terra, the mid-tier, fell 20%, to $2.00 and $12.00 from $2.50 and $15.00. Sol, then the flagship, stayed at $5.00 and $30.00 that day and gained Fast mode, up to 2.5 times the speed at twice the standard price.
Altman wrote the same afternoon that these were “major price cuts.” Cached Luna input dropped to $0.02 per million tokens. ChatGPT and Codex monthly fees and quota budgets did not change, though Terra and Luna started consuming fewer credits inside those plans. AWS Bedrock pricing began to follow later that day.
We are committed to pushing the model frontier across cost efficiency, capability, and speed.
Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API.
Luna and Terra’s lower prices are… pic.twitter.com/rFhK7XKedp
— OpenAI (@OpenAI) July 30, 2026
OpenAI credited its own stack. It said GPT-5.6 Sol had rewritten production GPU kernels and lifted speculative decoding, cutting serving cost about 20% and raising token-generation efficiency more than 15%. The company framed the discount as those gains passed through, with extra pressure from cheaper Chinese open-weight models sitting in the background.
THE PRICE CUT CALENDAR
- June 8, 2026: OpenAI files a confidential draft S-1 and says timing for a listing is still open.
- June 9, 2026: Anthropic ships Claude Fable 5 and Claude Mythos 5 at $10 input and $50 output per million tokens.
- June 11, 2026: People familiar with OpenAI’s plans describe a review of token prices aimed at developers and businesses.
- July 9, 2026: The GPT-5.6 family reaches general availability at launch rates, including Luna at $1.00 and $6.00.
- July 30, 2026: Luna falls 80% and Terra 20%. Sol keeps $5.00 and $30.00 and adds Fast mode.
- August 21, 2026: Sol moves to promotional pricing of $4.00 and $20.00, at least through November 21, 2026.
- September 2026: CFO Sarah Friar says the Luna cut helped drive roughly tenfold more use, and Altman later rules out a 2026 IPO.
The June leak was about whether OpenAI would blink. By late July it had blinked on the volume tier, and it had done so without touching Sol’s sticker that week.
How an 80% Cut Produced Tenfold Luna Traffic
Friar, OpenAI’s chief financial officer, took that result to Goldman Sachs’ Communacopia + Technology Conference in San Francisco in September 2026. She said the Luna cut helped drive a roughly tenfold increase in usage. Enterprise revenue rose 32% from June to July, against about 20% growth in overall annualized revenue in the same stretch. Consumer and enterprise were already a roughly even split by midyear, ahead of a year-end target.
If usage tracks billable tokens, an 80% unit cut against tenfold traffic is about twice the prior Luna spend, not a smaller check. That is the second-order arithmetic inside the “cheaper AI” headline. OpenAI still booked more enterprise revenue in July, which is what a price-elastic product is supposed to do for the seller.
THE LUNA VOLUME SNAPSHOT
- Unit cut: Luna’s standard API rate fell 80% on both input and output on July 30.
- Traffic: Friar said usage then rose roughly tenfold.
- Implied spend: Ten times the tokens at one-fifth the rate is about 2 times the old Luna bill if the mix holds.
- Codex: Friar put OpenAI’s coding tool at 25 million users in the same remarks.
Developers who live on the cheap tier did not treat the cut as spare budget. They routed more agent loops, more background jobs, and more first-pass drafts through Luna. Weekly caps on ChatGPT Plus still snap, so the constraint moved from dollars toward rate limits. A 5-hour throttle is the same fight in another costume: the meter is slower, the appetite is not.
GPT-5.6 Luna is the closest we’ve come to intelligence too cheap to meter. I’ve never seen a model this affordable be this powerful, it’s unlocking use cases for Replit we didn’t expect to build for a long time.
Michele Catasta, President and Head of AI, Replit, quoted on OpenAI’s July 30 product post
OpenAI’s own pitch is that Luna now does work that used to need a dearer model. It says Luna matches last year’s frontier class at about 6 cents on the dollar per task and at nearly nine times the speed. That is the company’s comparison, not an independent audit, and it is how the firm wants buyers to think: cost per completed job, not cost per token.
Anthropic’s Fable 5 Stayed at $10 and $50
Anthropic did not copy the Luna markdown. Claude Fable 5 launched on June 9 at $10 per million input tokens and $50 per million output tokens, which Anthropic called less than half Mythos Preview pricing. That list price was still the official rate in mid-September. The public Fable 5 with Mythos guardrails is the generally available twin of Mythos 5, which stayed gated for trusted cyber and science programs.
Fable 5 went dark on June 12 after a safeguard problem, then returned on July 1. Anthropic said classifiers would fire, on average, in fewer than 5% of sessions. The product OpenAI was answering in June was a dear, capable coding model, not a 20-cent workhorse. OpenAI’s counter was a split book: Luna and Terra for volume, Sol and later GPT-6 Astra for the top of the stack.
STICKER PRICES PER MILLION TOKENS
| Model | Input | Output | Status |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | Permanent cut, July 30 |
| GPT-5.6 Terra | $2.00 | $12.00 | Permanent cut, July 30 |
| GPT-5.6 Sol | $4.00 | $20.00 | Promo, at least through Nov. 21 |
| GPT-6 Astra | $10.00 | $50.00 | Frontier list, short context |
| Claude Fable 5 | $10.00 | $50.00 | Unchanged list since June 9 |
Those are the published rates per million tokens on OpenAI’s API page, with Sol still marked promotional. Astra’s short-context sticker matches Fable 5 on the nose, which is the point of the GPT-6 Astra cyber model launch sitting above Sol rather than undercutting Anthropic’s premium line. The price war the June stories promised is real on the bottom shelf and polite at the top.
OpenAI’s last private mark in March 2026 was about $852 billion. Anthropic’s Series H on May 28 landed at about $965 billion. Both labs had confidential IPO paper in motion in June. Cutting the cheap model is a volume bid in that climate. Leaving Fable 5 at $10 and $50 is a bet that coding and long-horizon work will still pay a premium.
Uber Had Already Spent the Year’s AI Budget
The June leak did not invent the bill shock. Uber said its 2026 AI coding budget was gone by April. It then put staff on spend tiers, with a base of $1,500 per employee per month per tool and a path to ask for more. That sequence landed before Luna’s markdown, which is why a cheaper token did not automatically read as a smaller Uber invoice.
Finance teams spent the spring shutting down the culture that produced those invoices. Internal leaderboards that ranked who burned the most tokens looked like adoption in January and like waste by May.
WHERE THE SPEND GOT CAPPED
- Uber: Full-year AI coding money was exhausted by April 2026, then capped at $1,500 per employee per month per tool.
- Amazon: An internal token leaderboard came down, with staff told not to use AI just to use AI.
- OpenAI itself: The company said the median researcher in its research group was using more than $600 in tokens a day by mid-August, up from $162 in July, with the 90th percentile above $7,000 a day.
- Altman’s benchmark: He said OpenAI’s top token user burns about 100 billion tokens a month, against about 100,000 a month six and a half years earlier.
Per-token prices had already been falling for years while agent tools multiplied the tokens behind each task. A linear chat in 2023 was a cheap round trip. An agent that plans, calls tools, retries, and writes tests can spend thirty times that on one job. Cheaper Luna makes those loops affordable, which is why they multiply.
Tokenmaxxing, the habit of using as many tokens as possible to look busy, ran into that math. The people who still want unlimited run-rate are the ones building products on the API. The people who want a lid are the CFOs who watched April land in February.
What ChatGPT Subscriptions Still Cost
The June headlines mixed ChatGPT the app with tokens on the API. OpenAI split those products on July 30. Subscription prices and quota budgets stayed put. What changed is how Terra and Luna count against those budgets, so a Plus or Business seat can run more of the cheap models before it hits a wall.
The August 21 Sol promo followed the same split. API and credit pricing on Sol dropped more than 20% for three months, with output 33% cheaper. OpenAI said Pro, Plus, and Business subscription usage stayed unchanged. If you pay a monthly ChatGPT fee, the Sol markdown is not a lower card charge. If you pay the API, it is.
That is why “ChatGPT prices” was always a muddy phrase. Seats are a cap plus a model picker. Tokens are a meter. OpenAI cut the meter on Luna and Terra, then discounted Sol’s meter for a defined window, and left the seat prices where they were. Anyone hoping the $20 or $200 plan would get cheaper from the June leak did not get that deal.
Free and Go users were pointed at Terra. Plus, Pro, Business, and Enterprise seats can pick Terra and Luna. The cheap model became the default workhorse in more places, including Auto-review in the ChatGPT app and the Codex CLI, which OpenAI said should cost about 10 times less after the Luna rate and the model upgrade from GPT-5.4.
OpenAI Is Testing Prices Tied to Outcomes
Friar also said OpenAI is trying prices tied to business results rather than tokens burned. That is the logical next product if CFOs will no longer sign a blank meter. A cost-per-resolution chat bot already sells that way in customer support. Coding and research are harder to wrap in a single outcome, which is why the test is a test.
If you’re deploying Luna and compare that to (Z.ai’s) GLM 5.3, for example, on a cloud layer, we are cheaper.
Sarah Friar, CFO, OpenAI, at Goldman Sachs’ Communacopia + Technology Conference
She used that line to argue that a hosted Luna call can undercut a Chinese open-weight model once you add the cloud bill. The competitive set in June was Anthropic. By September it was also GLM 5.3 on someone else’s GPUs. OpenAI’s answer is to make the low tier cheaper than the substitute, then sell industry stacks on top in chip design, life sciences, and financial services.
Altman, in an interview published on September 13, said a 2026 listing would be an “ill-advised moment” and put safety work ahead of going public this year. The confidential filing from June 8 is still paper, not a ticker. The company can keep cutting the volume tier, promoting Sol into November, and quoting cost per task without printing a quarterly gross margin for the public market.
Cheaper tokens arrived, as the June leak said they might. They did not retire the invoice. They changed what was worth automating, and they gave OpenAI a tenfold Luna spike to show for it.
Frequently Asked Questions
What is a token in OpenAI API billing?
A token is a billing slice of text, often a short piece of a word, counted separately for what you send (input) and what the model writes (output). Cached input is cheaper than a fresh prompt, so Luna cached input is $0.02 per million tokens against $0.20 for uncached short-context input, and OpenAI adds a 10% uplift on eligible regional processing endpoints released on or after March 5, 2026.
When does GPT-5.6 Sol promotional pricing end?
OpenAI says the Sol promo, $4.00 input and $20.00 output per million tokens on short context, is available at least through November 21, 2026, and it has not published the rate that follows. Long-context Sol under the same promo lists at $8.00 input and $30.00 output per million tokens, and Fast mode, Batch, and Flex inherit the discount.
How does GPT-5.6 Luna compare with Claude Fable 5 on list price?
Fable 5’s official API rate is $10.00 input and $50.00 output per million tokens, which is 50 times Luna’s $0.20 input rate and about 42 times Luna’s $1.20 output rate. That gap is why OpenAI can call Luna a high-volume stand-in while Anthropic keeps Fable 5 in the premium coding and long-horizon band.
What is Fast mode on GPT-5.6 Sol?
Fast mode replaced Priority Processing on July 30, 2026, and OpenAI says it delivers up to 2.5 times the speed of Standard processing at twice Standard’s price with no change in measured intelligence. Requests that still send service_tier as priority are billed and routed as Fast mode, so old client code does not have to be rewritten to keep the faster lane.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
