156 Math Steps a Byte, and Your Chatbot Uses One
A chatbot chip can do far more arithmetic than it ever gets to, because for every word it writes it must haul the whole model in from memory. Here is why that sets the speed of every answer, and what Nvidia and Groq are doing about it.
The speed of an AI answer is set by how fast a chip can read its own memory, not by how fast it can calculate, and that is why the most expensive hardware in the world spends most of its time waiting.
That sounds backwards. Chips are sold on their calculating power, measured in teraflops, meaning trillions of arithmetic operations a second. Yet when a language model writes a reply to you alone, nearly all of that power goes unused. The bottleneck has a name, the memory wall, and once you see it, most of the odd behaviour of AI services stops being odd.
What is the model actually doing when it writes?
It is guessing the next token, over and over. A token is the basic unit of text a model reads: sometimes a whole word, sometimes a piece of one, sometimes a bit of punctuation [4].
A model builds its answer one token at a time, and each new token depends on every token that came before it, including yours [4]. Engineers call this autoregressive generation.
That chain cannot be skipped ahead. The model cannot write the tenth word until it has the ninth.
Inference, the work of running a finished model, therefore has two phases with opposite problems. The first, called prefill, takes in your whole prompt in one parallel pass, and it is limited by how much the chip can compute [3][4][6]. The second, decode, produces the answer token by token, and it is limited by memory bandwidth, the rate at which data can be moved out of memory [3][4][6].
Prefill is a gulp. Decode is a drip.
Why does one word at a time keep the chip waiting?
Because every single token needs the whole model. A model is made of weights, billions of numbers (parameters) learned in training. To produce one token, the chip must read every weight out of its high-bandwidth memory, or HBM, the stacked memory sitting right beside the processor, and feed it to the compute units [4].
For a model with P parameters stored in 16-bit precision, that is 2P bytes read to do roughly 2P operations of arithmetic [4]. One multiplication per weight loaded. The weight is used once and thrown away, then loaded again for the next token.
The measure engineers use is arithmetic intensity, the amount of arithmetic done per byte fetched. In decode it is about one operation per byte [4]. An NVIDIA A100, a widely used data-centre GPU, can do 312 trillion operations a second but move only 2.0 terabytes a second, so it is built to do 156 operations for every byte it receives [4].
Decode asks for one. By that arithmetic the chip runs at well under one percent of its calculating capacity, and the rest of the time it waits for weights to arrive [4].
Here is the version for a six-year-old. Picture a chef who can chop a thousand vegetables a minute, but the only pantry is down a long hall, and she can carry back one small basket at a time. She is not slow at chopping. She is slow at fetching. Buying her faster knives changes nothing.
The picture fails in one place: a real chef can see what is in the basket, while the chip must fetch the whole pantry again for every single word.
How fast can it go, then?
There is a formula, and it is short. The ceiling on single-stream speed is memory bandwidth divided by the size of the model in bytes [3]. A one GB model on a device with 90 GB/s of bandwidth cannot exceed about 90 tokens a second, however many teraflops the marketing claims [3]. Real systems reach around 60 to 80 percent of that ceiling once cache reads, attention calculations and software overhead are counted [3].
People can feel these numbers. About 10 tokens a second is a comfortable reading pace, anything under 20 feels sluggish in a chat, around 40 starts to feel snappy, and 80 or more feels fast, which is the territory of hosted services such as ChatGPT [3].
The formula also produces some surprises. A flagship smartphone, at roughly 85 GB/s, and a desktop PC with dual-channel DDR5 memory, at roughly 90 GB/s, sit at almost the same point [3].
DDR5 is simply the standard memory sticks in an ordinary desktop PC, and dual-channel means two of them working side by side. On the one number that decides token speed, the phone in your pocket matches the tower under the desk [3].
Capacity is a separate matter. An RTX 4060 graphics card with eight GB of memory and a Raspberry Pi with eight GB hold the same amount, but the card moves data at 272 GB/s and the Pi at 17 GB/s, a 16-fold gap at the same size [3]. How much memory a machine holds tells you whether the model fits. How fast it can read tells you whether you will enjoy using it.
What is the KV cache, and why does it make things worse?
When the model writes a new token, it has to look back at everything already said. This look-back is called attention: the part of the model that works out which earlier words matter most for the next one. Without help, the model would redo that work for the whole conversation every time, and the effort would grow with the square of the conversation length [4].
So it keeps a stored record of what it worked out for each earlier token, a pair of items called a key and a value. That stored record is the KV cache [1][2][4]. The model reads the cache instead of recomputing, a swap of memory for speed [1][2][4].
One source compares it to a receipt that gets longer with every item you add [1]. Each new word adds a line, and for every new word the model reads the whole receipt again.
The cost is easy to count. The cache needs memory equal to the number of layers in the model, times the number of conversations handled at once, times the length of the text, times the bytes used per number, times two for keys and values, times two further internal sizes of the model, called key-value heads and head dimension, which you can safely skip [1]. Every one of those factors is linear, so doubling any one doubles the memory [1].
For a seven-billion-parameter model with 32 layers, at 2,048 tokens in 16-bit precision, the cache alone takes roughly one GB [1][4].
By simple arithmetic from those figures, the weights of that same model would take about 14 GB at 16-bit, so the cache is a modest add-on at that length. Stretch the conversation, paste in a long document or run a coding session with hundreds of lines of context, and it keeps growing, competing with the weights for the same scarce high-bandwidth memory [1][4].
It also competes for the same pipe. At the batch sizes typical of a single user, the cache reads in decode add to the bandwidth burden rather than relieving it, because the chip must now read the weights and the growing cache for every token [1][6]. That is why a long chat can feel slower than a short one.
Then why do AI services feel fast at all?
Because the providers serve many people at once. Batching means processing several users' sequences together, so that each weight, once fetched, is used for several tokens instead of one [3][4]. The chef carries the same basket back, but now she is chopping for ten orders.
That lifts the arithmetic per byte and moves the work away from the memory-bound regime towards the compute-bound one [3][4].
It does not help you personally. Batching does nothing for the latency of any single request, and every extra sequence needs its own KV cache, which fills the memory quickly [3][4]. In production, KV cache growth under load is a leading way for services to fail: more concurrent users, more caches fighting for the same memory, then latency spikes, lower throughput and fragmented memory, even while the compute cores sit underused [2][6].
That is the answer to the six-year-old's question. The machine can look like a genius at doing lots of work, because it serves a crowd and takes in a long prompt in one gulp. But your one reply still has to go word by word, and every word is a trip down the hall.
Can the wall be moved?
There are two ways. One is more bandwidth. The H100 offers 3.35 TB/s with HBM3, the B200 reaches eight TB/s with HBM3e, and the forthcoming Vera Rubin is put at roughly 22 TB/s with HBM4 [3][5].
The trouble is that compute has risen faster still, so the gap between what a chip can calculate and what memory can feed it has widened, not narrowed [3][5].
The wall is also old. For decades processors have outrun memory [2]. To see why language models hit it so hard, start with the arithmetic underneath. A matrix is just a grid of numbers, and most of AI is multiplying grids.
Earlier workloads such as image classifiers multiplied one big grid by another big grid, so every number fetched from memory was reused many times and the compute cores stayed fed. Decode breaks that.
It multiplies one short list of numbers, the single token being processed, by the model's grid of weights. Engineers call that a matrix-by-vector product. It loads the entire model to do almost nothing with each weight.
The other way is a different chip. Groq's Language Processing Unit, or LPU, holds the model's weights in SRAM on the chip itself rather than in off-chip HBM [5]. SRAM is the very fast memory normally used for small caches.
One LPU reaches roughly 80 TB/s of internal bandwidth, about 24 times the H100's 3.35 TB/s, which for models that fit removes the wall [5].
The price is capacity. SRAM is small next to HBM, so serving a large model takes many LPUs working together, and the cost of a deployment grows with model size in a way GPU deployments do not [5].
The industry is hedging. In December 2025 Nvidia agreed to license Groq's inference technology, in a deal reported at roughly $20 billion [5].
At GTC 2026 Nvidia showed the Groq 3 LPU, an SRAM-based decode co-processor placed inside the Vera Rubin platform, with Rubin GPUs doing prefill and the LPU doing decode [5]. The split follows the two phases above: GPUs for the gulp, the new chip for the drip.
What can be done with the chips we have?
Two tricks are common. KV cache quantization stores the keys and values in eight-bit or four-bit numbers instead of 16, cutting the cache by two or four times with minimal quality loss [2][4]. Speculative decoding has a small draft model guess several tokens ahead, then has the large model check them all in a single pass, turning idle compute into extra tokens [2][4].
Both work inside the constraint rather than removing it. So do batching and the SRAM chips. Whatever the design, the limit is how fast bytes travel from memory to the processor [1][3][4].
In my view that makes the wall a fact about the way these models write, one token after another, and not a mistake engineers forgot to fix.
What the sources leave open
The reporting here is technical writing from researchers, vendors and analysts, and it does not say how far the Groq 3 LPU has been deployed within Vera Rubin or what latency it has posted. It also gives no figures on how widely speculative decoding is used by the big chatbot services, or on how the memory wall feeds into what providers charge per token.
What to watch
The test comes when Vera Rubin and its LPU co-processor reach real customers. If splitting the work between two kinds of chip delivers fast single replies at a bearable cost, then the next few years of AI speed will be decided in memory design, not in teraflops. Watch the tokens-per-second figures, not the teraflops.
Every edition in brief, three times a day, on our Telegram channel, on Bluesky and on Threads.
Spotted an error? Tell the editors
- KV Cache Memory for LLM Inference - Interactive mbrenndoerfer.com
- What Is KV Cache? How It Works, Why It Bottlenecks & How to Fix It - WEKA weka.io
- The Memory Bandwidth Ladder: What Actually Decides How Fast Your LLM Runs mlechner.substack.com
- The Need for Speed: Why LLMs Generate Tokens Slowly - Conscious Engines consciousengines.com
- Groq: The LPU Architecture, the Nvidia Licensing Deal and the Neocloud Pivot | InferenceChips.com inferencechips.com
- What Actually Limits LLM Inference Speed? (GPU vs Memory vs KV Cache Explained) | Yotta Labs yottalabs.ai




