Connect with us

AI

OpenAI’s Cheaper Inference Is Feeding a $21 Billion Chip Bet

OpenAI cut guest ChatGPT inference costs by more than half, while Etched shipped Jane Street a rack and reached a $21 billion valuation.

Published

on

OpenAI engineers told colleagues in June they had cut AI inference costs by more than half on models they optimized. Logged-out ChatGPT traffic then ran on a few hundred Nvidia GPUs, a software win that used no new chips.

The same summer, Etched shipped its first inference rack to Jane Street and raised $700 million at a $21 billion valuation. Cheaper replies did not empty the order book. They showed that the bill that matters now is the cost of thinking, which is why a trading firm bought silicon instead of waiting for another software patch.

Guest ChatGPT Now Runs on a Few Hundred GPUs

The June work was a squeeze on servers OpenAI already owned. Engineers described the gain as a cut of more than half in the cost of running existing models, then applied it to visitors who use ChatGPT without an account. After that change, that slice of traffic sat on a few hundred Nvidia GPUs.

OPENAI’S JUNE INFERENCE SNAPSHOT

  • The cut: Engineers said the cost of running the models they touched fell by more than half, from software changes rather than new hardware.
  • Guest traffic: Logged-out ChatGPT then ran on a few hundred Nvidia GPUs, on a tier that offers only a limited set of features.
  • API margin: Internal figures put gross margin at 39 percent at the end of the first quarter of 2026, up from 33 percent a year earlier.
  • The target: OpenAI is aiming for 52 percent by December 2026, which is still a goal, not a reported result.

OpenAI has not confirmed the work in public, and it is not clear the same tricks carry over to paid ChatGPT or to the full API. Guest users already see a thin product, so a smaller GPU pile there does not prove the expensive reasoning models suddenly got cheap. Anthropic treats the same lever as a secret, with chief executive Dario Amodei calling the firm’s methods Compute Multipliers and keeping the details in-house so rivals cannot copy them.

Eight Months From $5 Billion to $21 Billion

Etched, a San Jose company founded in 2022, spent June 30 disclosing silicon, a backlog, and money it had raised in quiet rounds. Twenty-six days after a July step-up, Jane Street led another round that more than doubled the price tag again.

ETCHED’S 2026 FUNDING LADDER

Date What closed Capital Post-money
December 2025 Stripes-led round, later disclosed $500 million $5 billion
June 30, 2026 Stealth; working chip and backlog disclosed $800 million raised to date $5 billion
July 23, 2026 Sequoia-led round, with Andreessen Horowitz, Jane Street, and SK Hynix $300 million $10.3 billion
August 18, 2026 Jane Street-led round and first customer delivery $700 million $21 billion

Count the new checks after the June disclosure and the company has taken in $1.8 billion. Co-founder and chief executive Gavin Uberti said the industry still lacked hardware that could serve frontier models in a way that pays. Co-founder Rob Wachen put the operating idea in four words: production is the product.

The June release listed $1 billion in signed contracts for rack-scale systems, first-pass A0 silicon on TSMC’s N4P process, and a team of more than 400 people drawn mostly from Nvidia, Broadcom, Google’s TPU program, SK Hynix, and trading firms. Early racks were already running DeepSeek, Qwen, Mamba, and Llama, the company said, and were built for models of many shapes and sizes.

Cheaper Tokens Buy More Thinking Time

A guest-tier software cut looks like a smaller Nvidia bill. The demand side of the same summer says the opposite. Goldman Sachs Research, in a May 2026 note on the agent economy, forecast that global token use will rise 24-fold by 2030, to 120 quadrillion tokens a month, while the cost of each token keeps falling 60 to 70 percent a year. Agent jobs, the bank said, can burn 10 to 50 times as many tokens as a plain chat because they loop, call tools, retry, and sometimes run overnight.

That is why a lab can cut the cost of a logged-out reply and still need more specialized silicon. Reasoning models already spend far more compute per answer than a short chat, and agents multiply that again. Savings on the cheap tier free GPUs that get pointed at longer thoughts, not returned to the vendor. Goldman also lifted its 2030 forecast for AI-linked servers to about $1.24 trillion, which is a hardware number sitting on top of a cheaper token.

Wachen has described the gap in simpler terms: the cost of producing intelligence is now far below the value of that intelligence, which points to a long shortage of tokens. Etched’s own July note said under 1 percent of the world can reach the most advanced models, and that spreading them will take many gigawatts of better tokens-per-watt hardware.

Jane Street Took the First Rack Then Wrote the Check

On August 18 Etched said it had shipped its first rack to Jane Street and raised $700 million at $21 billion, with Jane Street leading after it tested the hardware. Kleiner Perkins, Sequoia Capital, Andreessen Horowitz, Peter Thiel, Tiger Global, Bain Capital Ventures, Blackstone, and earlier backers joined. The named customer and the lead investor are the same firm.

We tested the chip and are pleased with the early results. Etched’s unique approach to inference delivers the precision we will need to support our most demanding workloads. We’re excited to now have our own rack running in our datacenter.

Jane Street, in Etched’s August 18, 2026 announcement

A quant fund lives on latency and on being right in small increments, which is a harsher first trial than a demo day. It is also a narrow one. Jane Street is one rack in one datacenter, and the workloads, prices, and uptime have not been published. When the lead investor is also customer number one, diligence and a purchase order can blur, which is useful proof and a concentrated risk at the same time.

The June stealth post is the company’s own account of the silicon, the backlog, and the first racks.

How Etched Tries to Beat a Hot GPU

Etched is not selling a general-purpose GPU. It sells what it calls frontier inference clusters, full racks that pair purpose-built chips with boards, cold plates, interconnects, and software aimed at prefill (reading the prompt) and decode (writing the next token). The June technical note hangs on two bets that try to dodge the heat and memory traps that cap GPU inference.

THE TWO HARDWARE BETS

  • Low-voltage math: The company says its low-voltage inference architecture runs math blocks at under half the voltage of most AI chips, which lets it hold 80 percent or more of peak FLOPs on trillion-parameter sparse mixture-of-experts models without thermal throttling. Ordinary AI chips, it says, often settle under half of peak once they get hot. It points to bitcoin miners, which already run at under three times the voltage of typical AI chips.
  • Shared cluster memory: Cluster-scale memory is a low-latency pool across the scale-up domain, using a custom interconnect and an HBM plus SRAM mix so decode does not wait on a deep memory stack or a network hop between experts.
  • The whole rack: Etched builds chips, packages, boards, and cooling together, and says it opened a Taiwan factory and a 2 megawatt datacenter in its San Jose office so design and production sit in one loop.

Andrej Karpathy, who is on the cap table, said he was struck by the tokens per watt at interactive speed, and by the idea that this work sits at the opposite extreme from power lines: very low voltage and high current over tiny distances. That is an investor who has used the hardware talking about watts, not a slogan about replacing Nvidia.

The old transformer-only pitch is no longer the one Etched leads with. The June materials say the systems run Mamba as well as Llama-class models and are meant for many-trillion-parameter mixtures, long context, and agents. That walk-back matters, because a chip that can run only one recipe dies if the recipe changes. Flexibility costs efficiency. Etched is trying to keep both, which is the hard part, and the part that still lacks a public benchmark against a Blackwell rack.

Nvidia Still Runs Training as Inference Splits

None of this knocks Nvidia off training. CUDA and a general-purpose GPU remain the default for labs that train, retune, and serve mixed jobs on the same floor. Market trackers still put Nvidia near 80 percent of data-center AI accelerator sales in 2026. The fight is over the growing inference slice, where the job is repetitive enough that a custom chip can beat a GPU on cost per token, watts, and latency.

TrendForce expects custom AI chip shipments to grow 44.6 percent in 2026, against 16.1 percent for merchant GPUs, and it puts ASIC-based AI servers at 27.8 percent of the AI server market this year. Hyperscalers are already on that path with Google TPUs, Amazon Inferentia and Trainium, and in-house designs at Meta and Microsoft. Nvidia’s own answer has been to absorb specialists: it licensed Groq’s inference design in a deal valued at about $20 billion in December 2025, rather than leaving that architecture as a pure rival.

So the second-order read is not that Nvidia’s training franchise is over. It is that a software cut at OpenAI and a $21 billion tag on Etched can both be true because the token pile is still growing. Share can slip on inference while the dollars in that slice rise. Custom silicon, merchant GPUs, and a handful of startups can all book more work in the same year.

A $1 Billion Backlog Still Has to Ship

Etched’s own August note is modest on the next problems: new factories, global supply, fleet software, and kernel agents that improve themselves. A path to gigawatt scale in 2027 is a plan. One rack at Jane Street is a start. Signed contracts are not the same thing as racks that stay up, yield that holds, and software that keeps pace when models change.

WHAT WE KNOW

  • The round: Jane Street led $700 million at $21 billion on August 18, 2026, after taking the first rack.
  • The backlog: Etched disclosed more than $1 billion in signed customer contracts on June 30, 2026, and said production had started.
  • The software cut: OpenAI engineers described a more-than-half drop in inference cost on models they optimized, with logged-out ChatGPT then running on a few hundred GPUs.

WHAT IS UNCONFIRMED

  • OpenAI’s own voice: The company has not published the method, the before-and-after GPU counts, or whether the gain extends past the guest tier.
  • Independent speed tests: Etched has not released public numbers that pit a production rack against Nvidia hardware on a shared model and batch size.
  • How firm the orders are: Outside the Jane Street delivery, contract terms, cancelation rights, and the number of racks in the field have not been disclosed.

The summer leaves a simple pile of facts. OpenAI found room in software to run a cheap ChatGPT tier on a few hundred GPUs. Etched, priced at $21 billion, has one named rack in a customer’s floor and a billion-dollar backlog still to build. If those contracts turn into fleets, the inference bill moves toward purpose-built silicon even as each token gets cheaper. If they do not, the $21 billion figure will look like a purchase order that never became a factory.

Disclaimer: This article is news reporting and analysis for general information only. It is not investment, financial, or trading advice, and it does not recommend buying or selling any security, private-company stake, or other financial product. Anyone making a money decision should speak with a licensed financial adviser who can review their own situation. Round sizes, valuations, margins, and product status come from company statements and the research cited here and may change.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending