OpenAI Has Not Said Why Its Fastest Tier Changed Chips. Memory Is the Clue
OpenAI's premium speed tier ran on Cerebras chips for one model and runs on Nvidia GPUs for the next. Behind the switch is how an AI answer gets made, and why the wait is mostly about memory, not arithmetic.
OpenAI's premium Ultrafast speed tier, launched in August on Cerebras chips, now runs on Nvidia GPUs for its newest model, GPT-6 Astra, and the switch shows how AI answers actually get made.
The move was announced at OpenAI's DevDay event on September 29, 2026. Nvidia confirmed on its own blog that Astra Ultrafast runs on its Blackwell chips [1]. A report from the research firm SemiAnalysis said a second model, GPT-6.1 Sol Ultrafast, would also run on Nvidia hardware [3][6]. Shares in Cerebras fell from $206.75 on September 25 to $166.43 on October 2, a drop of nearly 20 percent, though a post-IPO share unlock and insider-sale filings landed in the same window [3][6].
So why would OpenAI leave the chip that gave it 750 tokens a second? Neither company has explained. What the sources do show is how the machinery works, and that is the best guide to what such a move could mean. To get there, start with what happens after you press enter.
What happens between pressing enter and seeing the answer?
The step is called inference: a trained model takes your prompt and writes a reply. It writes one token at a time, where a token is a chunk of text, often a short word or part of one. Each new token depends on every token before it, so the model cannot skip ahead or write the whole reply at once.
For every token, the model consults its full set of learned weights, the numbers that hold what it knows. For the largest models these run to hundreds of billions of parameters. Speed, then, is the number of times per second the chip can finish that loop.
You might assume the limit is arithmetic, that a faster calculator means faster text. It is not. Hardly any time goes on the maths.
Why is the wait a memory problem?
On a conventional GPU, the weights sit in a separate memory bank beside the processor, called high-bandwidth memory, or HBM. For every token written, the weights must cross from that bank into the compute cores [2][5]. As models grow, inference is limited less by how fast the chip can calculate and more by how fast it can move data. Engineers call this the memory wall [2][5].
Picture a brilliant cook who can chop at lightning speed but whose ingredients are in a warehouse across the road. The knife is idle most of the time. Faster knife work changes nothing. A shorter walk changes everything.
The analogy stops at one point. The cook makes one trip per dish, while the chip makes the whole trip for every single word.
What is the KV cache, and why does it matter?
The KV cache is the second half of the puzzle. "KV" stands for key and value, two sets of numbers the model calculates for each token as it reads, which record what that token contributes to the meaning of the text [4].
The cache is a store of those numbers for every token already processed, so the model can look back at what came before instead of working it out again for each new word [4]. Think of a meeting where someone takes notes. Without them, every time a new point is raised, the group would have to replay the whole meeting from the start.
The saving is large. Without a cache, writing a 1,000-token reply would mean the 1,000th token recomputing the numbers for all 999 before it [4]. The cache trades memory for computing time, and it is what makes text appear fast enough to feel like a conversation [4].
The bill comes due in memory. The cache grows with the length of the conversation, the number of requests being served at once, and the size of the model. With context windows of tens or hundreds of thousands of tokens, the cache alone can take tens of gigabytes of GPU memory, competing with the weights for the same limited space [4].
Worse, while the model is generating, the cache is read again and again from the chip's memory. That makes memory bandwidth, how much data can move per second, the dominant bottleneck. At low batch sizes, speed is set by how fast a chip can read its own memory, not by its raw calculating power [4].
Batching, in this sense, means grouping several users' requests so the chip handles them together. A big batch squeezes the most total work out of the hardware. A low batch, one or a few requests at a time, favours how quickly you personally see words appear [6]. Ultrafast is a bet on the second.
How do Cerebras and Nvidia attack the wall?
They come at it from opposite ends.
Cerebras removes most of the walk. Each of its WSE-3 chips is an entire silicon wafer, packing 900,000 AI-optimised cores and 44 GB of fast on-chip memory called SRAM, with 21 petabytes per second of aggregate memory bandwidth [2][5]. The weights sit right beside the compute cores. Nothing has to be hauled in from across the road [2][5].
Set that beside one Blackwell GPU, which has 192 GB of HBM3e memory at eight terabytes per second [1]. On paper, the Cerebras wafer's bandwidth is more than 2,500 times higher. The GPU, though, holds more than four times as much memory per chip.
That gap is the trade-off. A big model does not fit on one wafer. So Cerebras splits it across several CS-3 systems, which are its full machines built around a wafer.
A model is a stack of layers, each one a stage of processing that passes its result to the next, and the split is made between those stages. Each wafer keeps its assigned layers in local memory, and only the intermediate results, called activations, pass between wafers [5]. It works, and it gave GPT-5.6 Sol up to 750 tokens a second with the same weights, precision and reasoning settings as the standard version [2][5].
Nvidia takes the other route: more bandwidth, more memory and smarter handling of what fills it. Blackwell offers a compressed format for the KV cache, called NVFP4, that cuts the cache's footprint by up to 50 percent [1].
NVFP4 is a way of storing each number in fewer bits, the ones and zeros a computer keeps, so the same cache takes up less room. Blackwell also lets the cache spill over into CPU memory at 900 GB/s [1].
And it offers programmability. Nvidia's software ecosystem lets engineers write their own low-level routines, called kernels, that tell the chip exactly how to do its work.
So why is Astra Ultrafast on Nvidia?
Neither OpenAI nor Cerebras has commented officially on the reported infrastructure change, and none of the sources lays out OpenAI's reasoning. What follows is what the record shows and the paper's own reading of it.
The record first. OpenAI says it used its own models to write high-performance kernels for Blackwell, and Nvidia says that work is what accelerates Astra Ultrafast [1]. Philippe Tillet, OpenAI's inference lead, said Nvidia's tooling and documentation let OpenAI make its models "exceptionally good at programming Blackwell and Rubin GPUs" [1]. Uday Ruddarraju, OpenAI's CTO of compute, credited Nvidia's programmability for the acceleration [1]. Because the software can be updated after deployment, OpenAI can keep testing and rolling out improvements [1].
These are statements in a vendor's blog, which naturally casts Nvidia well.
The numbers are different too. Astra Ultrafast reaches up to 300 tokens per second, and OpenAI describes it as up to eight times faster in Codex and up to six times faster in the API than standard mode [1]. GPT-5.6 Sol on Cerebras hit 750, which is 2.5 times as many.
But these are different models, so the figures are not a head-to-head test of chips [1][2][5][6]. Nvidia's blog gives no absolute speed beyond the "up to" claims.
The paper's reading is this. Ultrafast was never a fixed speed. It is a purchase of speed, and what OpenAI buys depends on the model it is selling. A newer model is likely larger and more demanding, and a Cerebras wafer has far less memory per chip than a GPU.
SemiAnalysis speculated that constraints tied to Cerebras's on-chip memory architecture could be a factor, but that is unconfirmed. If so, 300 tokens per second on a bigger model from a platform OpenAI can tune after launch may beat 750 on a platform that is harder to scale. We cannot say that is what happened.
Supply matters as well. A faster tier is only useful if there are enough chips to sell it at scale.
SemiAnalysis said the new tier was running on Nvidia GPUs at a low batch size, the regime where Cerebras's on-chip memory had held a distinct advantage [6]. If that holds, Nvidia has closed some of that gap.
There is a wrinkle. The SemiAnalysis report concerns GPT-6.1 Sol Ultrafast, a cheaper model that OpenAI says approaches Astra's intelligence at roughly one-fifth of Astra's token price [3][6]. It is promised as coming soon, and it is a different model from Astra [1].
Is this a divorce?
Not on the evidence so far. Sam Altman later called Cerebras a "close partner", and OpenAI has not cancelled a multiyear agreement reported to be worth more than $10 billion for 750 megawatts of compute from 2026 through 2028 [3].
Cerebras, meanwhile, reported a GAAP net loss of $450.5 million in its latest quarter [6]. GAAP is the standard accounting measure, which counts every cost, including ones that involve no cash. Most of the loss was tied to stock-based compensation, which means paying staff partly in company shares [6].
Insiders sold about $266 million in shares in the three months before the drop [6]. That helps explain why a single report moved the stock so hard.
So what does "Ultrafast" mean? It means a model served at a higher speed than standard, at a higher price, on whatever hardware its provider judges best. Ultrafast costs six times the standard API price, though the hardware economics behind that price are not public.
The label does not name a chip, a speed or a quality level. The August Cerebras version kept the same model and settings as the standard one, and the claim should be checked each time for what has been traded away [2][5].
Ask three things of any speed claim. How big is the model? How many users share the chip? What does the speed cost, per token and in chips?
The answers differ every time. The sources point to a portfolio of hardware matched to different models, workloads and prices, not a single winner.
What to watch
GPT-6 Astra Ultrafast is already available in the OpenAI API and to ChatGPT Work and Codex users on Pro 500 and Enterprise plans [1]. GPT-6.1 Sol Ultrafast is promised as coming soon, and where it runs will show how far the Nvidia shift goes [1]. Neither OpenAI nor Cerebras has yet commented on the SemiAnalysis report.
Every edition in brief, three times a day, on our Telegram channel, on Bluesky and on Threads.
Spotted an error? Tell the editors
- How NVIDIA GPUs Help Accelerate OpenAI's GPT-6 Astra Ultrafast | NVIDIA Blog blogs.nvidia.com
- GPT-5.6 Sol Ultrafast on Cerebras | Up to 750 Tokens/s cerebras.ai
- Cerebras Stock Drops 20% After OpenAI Turns to Nvidia Chips: What It Means for the AI Chip Race| KuCoin kucoin.com
- What Is KV Cache? How It Works & How to Optimize It | WEKA weka.io
- Cerebras Serves GPT-5.6 Sol at 750 Tokens/Second cerebras.ai
- Cerebras Tumbles 7% as Report Says OpenAI's New Ultrafast Model Runs on Nvidia Chips - BigGo Finance finance.biggo.com




