Connect with us

AI

OpenAI Ultrafast Puts GPT-5.6 Sol on Live Workloads

Ultrafast mode delivers GPT-5.6 Sol at 750 tokens per second on Cerebras, collapsing overnight research and outage loops into interactive sessions for early.

Published

on

OpenAI launched Ultrafast mode for GPT-5.6 Sol on August 13, 2026, a new API service tier that runs the flagship model up to 14× faster than Standard processing and generates up to 750 output tokens per second. Powered by Cerebras wafer-scale systems, the tier keeps full Sol intelligence while targeting workflows where every second counts.

The preview is limited to select customers. OpenAI is studying where the speed shift changes real products before wider rollout.

That study period matters as much as the headline rate. OpenAI is not treating the tier as a finished product line. It is treating it as a live experiment in which speed, capacity and product design have to move together before the company opens the gates wider.

The Numbers Behind the New Tier

Ultrafast is not a smaller or distilled model. It serves the same GPT-5.6 Sol that already leads on coding, knowledge work and science benchmarks, only at far higher output rates. OpenAI frames the goal as more useful work per second rather than another quality jump.

  • 750 output tokens per second peak generation rate on Cerebras hardware
  • Up to 14× faster than Sol Standard processing
  • 5.6× end-to-end speedup on GDP-Val knowledge-work tasks with no quality drop
  • 11 hours 11 minutes to finish Humanity’s Last Exam’s 2,500 questions versus more than three days for a leading rival model at comparable accuracy

Token costs for the tier remain undisclosed during the preview. Businesses join a waitlist by describing workload, latency needs and expected volume while OpenAI evaluates early results.

Read together, the figures describe two different kinds of gain. Peak token rate is the hardware story. End-to-end speedup on GDP-Val is the product story: the same answers, delivered fast enough that a full knowledge-work loop shrinks instead of stretching. The Humanity’s Last Exam timing makes the same point at exam scale. A job that once ran past three days now fits inside a single long work session at comparable accuracy.

Because costs stay hidden in preview, buyers cannot yet model unit economics. They can only describe the jobs where latency is the binding constraint and wait for OpenAI to decide which of those jobs justify scarce wafer-scale capacity.

Cerebras Keeps the Weights on the Chip

The speed comes from a hardware choice that sidesteps the usual GPU bottleneck. On conventional accelerators, large-model inference repeatedly shuttles weights between limited on-chip memory and off-chip storage. That data movement dominates latency once models grow.

Cerebras packs 44 GB of SRAM on each wafer-sized chip. Weights stay resident. Tokens stream through pipelined layers across wafers with single-cycle local access and explicit mesh routing between hundreds of thousands of processing elements. The company describes the architecture as purpose-built for frontier inference rather than a general-purpose GPU cluster.

  • Wafer-scale engine with roughly 900,000 cores
  • On-chip SRAM measured in tens of gigabytes per wafer
  • Aggregate memory bandwidth in the tens of petabytes per second
  • Dataflow programming model that routes messages instead of launching thread blocks

Andrew Feldman, Cerebras CEO and co-founder, said the combination proves speed and intelligence are no longer mutually exclusive. The partnership itself dates to a multi-year capacity deal announced earlier in 2026 that targets hundreds of megawatts of wafer-scale systems for OpenAI inference.

The design choice is simple to state and hard to copy on ordinary clusters. If weights never leave the wafer, the model does not pay a memory-tax on every layer. Pipelined tokens then move through local mesh routes instead of waiting on off-chip round trips. That is why Cerebras can claim frontier scores and interactive rates in the same hosted tier, and why OpenAI’s capacity deal is measured in megawatts rather than in a handful of racks.

Feldman’s line about speed and intelligence lands because the preview keeps full Sol weights in play. The tier does not win by shrinking the model. It wins by removing the shuttle that usually throttles a large one.

Early Customers Already Change Their Loops

OpenAI is testing the tier with companies in coding, commerce, financial research, support and interactive applications. Named early users include Jane Street, Podium, Basis and Rogo. Their feedback focuses less on raw throughput and more on what becomes practical when the model keeps pace with a human or a live event.

The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.

John Crepezzi of Jane Street’s AI Assistants group said that after early exposure. Courtland Lykins, Product Lead for Voice AI at Podium, reported that Ultrafast “completely changes the call experience for the more complex work.” Mitch Troyanovsky, co-founder at Basis, noted that the combination of intelligence and speed removes a barrier that previously blocked truly synchronous products. Alex Wang at Rogo said complex financial research starts to feel like a real-time interaction.

These are not marketing placeholders. They describe a shift from batch-style prompting to continuous collaboration with the model still in the conversation.

The named accounts span different surfaces, yet the pattern is shared:

  • Jane Street: developers stay in a focused loop beside the model instead of waiting on batch turns
  • Podium: complex voice work keeps conversational flow on live calls
  • Basis: synchronous products stop looking blocked by model lag
  • Rogo: financial research starts to feel like a live exchange

None of those quotes celebrate a leaderboard. They describe a change in when the model is allowed into the critical path. Once replies arrive at human pace, teams stop parking hard questions for later and start asking them while the event is still open.

How the Speed Compares with Rival Fast Modes

Anthropic already offers a fast mode for its Opus models that delivers up to 2.5x higher output tokens per second at premium pricing. Cerebras and OpenAI position Ultrafast further out on the curve. Using third-party output-speed figures, they report Sol Ultrafast running roughly 5× faster than Claude Opus 4.8 in Fast mode and 11× faster than Claude Fable 5.

Mode / Model Peak claim Relative note
GPT-5.6 Sol Ultrafast Up to 750 output tokens/sec Full Sol intelligence
GPT-5.6 Sol Standard Baseline Up to 14× slower than Ultrafast
Claude Opus Fast ~2.5× over its standard ~5× slower than Sol Ultrafast per Cerebras
Claude Fable 5 Standard frontier pace ~11× slower than Sol Ultrafast per Cerebras

Exact cross-lab comparisons always depend on workload, batching and measurement date. The directional gap is large enough that OpenAI can market a distinct speed class rather than a modest acceleration. Chinese open-weight releases such as Moonshot’s Kimi K3 continue to pressure closed labs on capability and local deployment, yet they do not yet match this combination of frontier scores and sub-second interactive rates in a hosted API.

Anthropic’s fast mode shows that premium latency is already a recognized product category. Ultrafast’s claim is that the category now has a much higher ceiling when the stack is built around wafer-scale SRAM instead of a tuned GPU path. Open-weight models keep pressure on price and local control. They still leave a gap for buyers who want frontier scores and hosted interactive rates in one API call.

What Changes When Frontier Models Keep Pace

The second-order effect is larger than any single benchmark. Until now, teams that needed top-tier reasoning usually accepted multi-second or multi-minute waits, or they switched to smaller models for anything interactive. Ultrafast collapses those choices for the highest-stakes moments.

  • Incident response: read logs, traces and engineer reports while an outage is still unfolding and help prepare a fix before the window closes
  • Financial and security monitoring: assess transactions or market signals as conditions change rather than after the fact
  • Customer support and voice: resolve multi-step issues without breaking conversational flow
  • Commerce: answer inventory and personalization questions while a shopper is still deciding
  • Research: turn overnight experiment batches into multiple daytime iteration cycles

OpenAI’s own developers already run the tier during incidents for exactly these steps. Research loops that once spanned a night now support several cycles in a single workday. That pattern matches what early external customers describe: the model stops being a side tool and becomes part of the critical path.

The same Sol model that drew attention during the GPT-5.6 family launch with workplace agents and later earlier Sol autonomy reports on file deletion now gains a latency profile that makes those agentic behaviors feel immediate rather than experimental.

The practical change is a collapsed menu of tradeoffs. Teams no longer have to pick intelligence or interactivity for the moments that matter most. They can keep Sol on the line while an outage, a trade, a call, or a shopping session is still live. Agentic habits that felt experimental at slower rates start to look like ordinary tooling once the wait drops away.

The Waitlist Filters Workloads That Matter

OpenAI is not opening Ultrafast as a default switch on every Sol call. Select customers join a waitlist by describing workload, latency needs and expected volume. That intake is a filter as much as a queue. It tells the company which jobs actually move when the model answers at up to 750 output tokens per second.

The preview domains already hint at the filter’s bias. Coding, commerce, financial research, support and interactive applications all share a trait: a human or a live market is waiting on the other side of the response. Batch report generation can tolerate Standard processing. A voice turn, a trading desk prompt, or an incident thread cannot.

Capacity on specialized silicon is finite. Hundreds of megawatts of wafer-scale systems sound large in a press release and still look tight once voice, trading and reliability workloads compete for the same resident weights. The waitlist lets OpenAI watch where the 5.6× end-to-end style of gain shows up in shipping products before it promises the tier more broadly.

Pricing remains undisclosed for the same reason. Until the company sees which loops truly change, it cannot set a durable premium against Standard processing or against rival fast modes. Early access is a measurement program with revenue attached later.

Speed Turns Agents From Demo to Default

Workplace agents and autonomy features already traveled with the GPT-5.6 Sol family. What they lacked was a latency profile that matched how people actually work. Multi-second pauses break trust in an agent loop. Multi-minute pauses push the same agent back into overnight batch jobs.

Ultrafast shortens that gap. When output can peak at 750 tokens per second and Standard is up to 14× slower, an agent that plans, calls tools and revises mid-task can stay inside a single human turn. The file-handling and office-agent behaviors that drew earlier attention stop reading like demos and start reading like default paths for high-stakes work.

Early customer language points the same way. Synchronous products, live call handling and real-time research are agent-shaped problems. They need the model present for several steps in a row, not for one slow completion at the end. Full Sol intelligence at Ultrafast rates is what makes that presence affordable in product design terms.

Internal incident use closes the circle. Engineers who keep judgment and deployment in human hands still benefit when hypothesis generation no longer inserts dead time between alerts. The agentic pattern is not a separate product. It is Sol staying fast enough to remain on the critical path.

OpenAI’s Internal Teams Already Live on Ultrafast

Inside the company a group of engineers uses Sol on Ultrafast to shorten the gap between an alert, a hypothesis and the next action. They stay responsible for judgment and deployment. The model simply removes the dead time that used to stretch every investigation.

Sachin Katti, OpenAI VP of Compute Strategy and GPT-Infra, said the company is exploring what becomes possible when customers get the intelligence of its most capable models with significantly lower latency, starting small and expanding as capacity and learnings allow. Pricing, broader availability and measured impact on product design remain open questions. Capacity on specialized silicon is finite, and demand from voice, trading and reliability workloads could outrun supply quickly.

For now the preview proves the technical claim. Frontier intelligence no longer has to wait. The workflows that form around that fact will decide how permanent the new speed class becomes.

The open questions Katti leaves on the table are the ones buyers will watch next: when the waitlist turns into general availability, how token pricing lands against Standard, and whether product teams redesign around continuous collaboration or merely sprinkle faster calls into old batch flows. The hardware proof is already public. The product proof will be written by the workflows that stick.

Logan Pierce is a writer and web publisher with over seven years of experience covering consumer technology. He has published work on independent tech blogs and freelance bylines covering Android devices, privacy focused software, and budget gadgets. Logan founded Oton Technology to publish clear, no nonsense tech news and reviews based on real hands on testing. He has personally tested and reviewed dozens of mid range and budget Android phones, written extensively about app privacy, and built and managed multiple WordPress publications over the past decade. Logan holds a bachelor's degree in English and studied digital marketing at a certificate level.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending