Articles

OpenAI Ultrafast: GPT-5.6 Sol at 14x Speed on Cerebras

OpenAI's new Ultrafast tier runs the full GPT-5.6 Sol model at up to 750 tokens per second — 14x faster — on Cerebras wafer-scale chips instead of GPUs.

Chisato Chisato · · 5 min read
A polished silicon wafer patterned with chip dies catching the light

The battle in AI has quietly shifted from how smart a model is to how fast it answers. On Wednesday, August 13, 2026, OpenAI previewed Ultrafast, a new service tier in its API that runs the full GPT-5.6 Sol model at up to 14 times the speed of standard processing — and it does so not on the GPUs that power almost all frontier inference, but on wafer-scale chips built by Cerebras Systems.

What OpenAI announced

Ultrafast is a new API tier, available today in a limited preview to a select group of customers, with OpenAI saying it will expand access “as capacity grows.” The pitch is simple and specific: it is the same GPT-5.6 Sol model, with the same intelligence as the standard version, running dramatically faster.

The headline figure is up to 750 output tokens per second — roughly the pace at which the model can generate text — compared with the tens-to-low-hundreds of tokens per second typical of frontier models on conventional hardware. OpenAI framed that as up to 14x faster than GPT-5.6 Sol on standard processing, with no change to the underlying model weights or reasoning quality.

That last point is the differentiator. The industry already has fast models, but they are usually smaller, distilled, or otherwise cut down to hit their speed. Ultrafast keeps the flagship large language model intact and instead attacks the hardware bottleneck that governs how quickly any given model can respond.

The Cerebras angle

The speed comes from Cerebras, which said separately that it is powering the Ultrafast tier. Cerebras builds the Wafer-Scale Engine, a processor the size of an entire silicon wafer rather than a fingernail-sized die. Its defining feature is on-chip memory: each wafer carries roughly 44 GB of SRAM directly on the silicon, keeping model weights next to the compute instead of shuttling them back and forth from external memory.

That design targets the single biggest constraint on inference speed. On conventional accelerators, generating each new token requires streaming the model’s weights out of high-bandwidth memory, and the gap between memory bandwidth and raw compute — not arithmetic — is what caps how fast a large model can talk. By holding weights on-chip, Cerebras sidesteps that ceiling, which is how it reaches 750 tokens per second on a full frontier model.

To dramatize the difference, Cerebras cited a benchmark run on Humanity’s Last Exam, a 2,500-question test spanning graduate-level subjects. GPT-5.6 Sol on Ultrafast, it said, answered the entire question set in just over 11 hours, versus more than three days for Claude Fable 5 running at conventional speeds — a rough illustration of what a 10-plus-times throughput advantage looks like when a model has thousands of hard problems to grind through.

Why speed, and why now

OpenAI positioned Ultrafast squarely at production use cases where latency, not intelligence, is the binding constraint. It named voice interfaces, customer service and support, developer agents, e-commerce, financial market analysis, and security incident response as the target workloads.

The common thread is that these are agentic and interactive applications. A voice assistant that pauses for three seconds feels broken; a coding agent that has to make dozens of sequential model calls to complete a task is bottlenecked by the slowest link in the chain. As AI shifts from single-shot chat toward multi-step agents that call a model repeatedly, per-token speed compounds: a workflow that makes 50 model calls in sequence finishes 14 times sooner if each call returns 14 times faster. For those workloads, throughput is not a nicety — it is the product.

There is also a competitive dimension. Fast-inference specialists — Cerebras among them, alongside rivals building custom silicon — have spent the past two years arguing that the future of AI economics runs through inference, not training, and that purpose-built chips can beat general-purpose GPUs on tokens-per-dollar and tokens-per-second for that job. OpenAI putting its flagship model on Cerebras hardware, even in preview, is a notable validation of that thesis from the most-watched lab in the industry.

The pricing subtext

OpenAI did not lead with price, but Ultrafast lands in the middle of an intensifying cost war. The company has spent 2026 cutting prices on GPT-5.6 to fend off rivals, and speed tiers are the natural next axis of competition once raw capability converges across the leading labs. A faster tier lets OpenAI capture latency-sensitive enterprise demand — the kind willing to pay a premium for responsiveness — without touching the price of its standard offering.

It also diversifies OpenAI’s compute supply. The company reportedly generates on the order of $2 billion a month in revenue while remaining unprofitable, and every workload it can move onto more efficient inference hardware eases the enormous compute bill behind those numbers. Spreading inference across more than one hardware vendor is both an economic hedge and a supply-chain one.

The caveats

Ultrafast is a preview, not a general release. Access is limited to select customers, and OpenAI was explicit that broader availability depends on how fast Cerebras capacity comes online — a meaningful constraint given that wafer-scale systems are produced in far smaller volumes than mainstream GPUs. The 750-tokens-per-second and 14x figures are also stated as up to numbers, which in practice vary with prompt length, output length, and load.

And the benchmark theater, while striking, should be read as a throughput demonstration rather than a claim about answer quality. OpenAI’s own framing is that intelligence is held constant; the entire value proposition is that you get the same GPT-5.6 Sol, only faster. Whether that speed holds at scale, and at what price, is what the preview period will reveal.

What it means

Ultrafast is a marker of where the AI race is heading. With the leading models clustering at similar capability levels, the frontier is moving from how good to how fast, how cheap, and how reliably at scale — and speed is the dimension enterprises can feel immediately in voice, support, and agent products. OpenAI is signaling that it intends to compete on that axis, not just on benchmark scores.

The clearest winner is Cerebras, which gets the most valuable possible reference customer for its wafer-scale architecture and a concrete, public demonstration that its on-chip-memory approach delivers something GPUs structurally cannot match on latency. For the broader ecosystem of custom inference silicon, OpenAI’s endorsement lends credibility to the argument that the inference layer is where specialized hardware finally breaks the GPU’s grip.

The pressure lands on rivals whose fastest offerings are smaller or distilled models: OpenAI is claiming full-flagship intelligence at specialist speed, collapsing a trade-off competitors have leaned on. It also lands, subtly, on the general-purpose accelerator incumbents, for whom inference is the fastest-growing and most contested part of the market.

What to watch next: how quickly Ultrafast moves from preview to general availability, what OpenAI charges for the tier, whether other labs answer with their own fast-inference partnerships, and whether the wave of latency-sensitive agent products now under construction standardizes on this kind of speed as table stakes. The intelligence race is not over — but for the first time, the clock is the headline number.

Chisato Chisato · · 6 min read

OpenAI Dots: Always-On ChatGPT Agents Explained

OpenAI's Dots are always-on agents with their own cloud computer, powered by GPT-6 Astra. How they work, who gets them, and the safety questions.

#AI #OpenAI #AI Agents