Mind the Gap: DeepSeek's Memory Squeeze Blurs Facts at Regular Intervals
A technique that shrinks the memory long-context AI models need makes them much better at finding some facts than others, depending on where in the text a fact falls. The paper tested four DeepSeek-V4 variants and small models trained from scratch.
Researchers at ByteDance Seed report that the memory-saving compression inside DeepSeek's latest models leaves a repeating blind spot, with retrieval accuracy swinging by 15 to 40 percentage points depending on where a fact sits.
The paper, "Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression," was posted to arXiv on September 28, 2026 as a pre-print. It has eight authors from ByteDance Seed, with core contributors also affiliated with Princeton, Stanford and UC Berkeley [1][4].
What the technique does
A language model keeps a running record of everything it has read, called the KV cache, and consults it to write each new word. At very long context lengths that record eats GPU memory and slows the model down.
Chunked compression shrinks it. A window of W tokens slides across the text in steps of S tokens, and each window is pooled into a single cache entry [1][2][4]. DeepSeek-V4 leans on this heavily. It has no full-attention layers at all, according to a technical survey by Julien Simon.
It reads the last 128 tokens exactly and reaches anything older only through compressed summaries. About half its layers pool windows of eight tokens with a stride of four, and the other half squeeze every 128 tokens into one entry [3].
The ByteDance team's point is that this creates a new coordinate for every token. They call it the token's phase: its position modulo the stride, meaning where it falls within each repeating step. Phase decides which neighbours a token is blended with, and therefore how legible it is later [1][2][4].
The 40-point swing
The team ran a needle-in-a-haystack test, hiding a target among 16,000 key-value records in 128,000 tokens of context. All four DeepSeek-V4 variants they tested (Flash-Base, Flash-0731, Pro-Base and Pro-0813) rose and fell in accuracy with a period of four, matching the model's stride of four. The gap between best and worst phase ran from 15 to 40 points [2][4].
The 40-point figure comes from base checkpoints, the models before post-training. Post-training raised accuracy and narrowed the gap but kept the periodic pattern [2][3][4].
DeepSeek-V4.1-Flash compresses with a stride of two, and its pattern has a period of two. Even-positioned targets scored 94.7 to 95.3 percent and odd-positioned ones 89.2 to 90.3 percent, a gap of roughly five to six points [2][3][4]. Simon's survey puts the best-to-worst gap at 6.09 points against a mean accuracy of 92.38 percent [3].
DeepSeek's own V4.1-Flash technical report, published September 17, 2026, had already conceded that "the newly introduced architectural changes also create robustness boundaries that have yet to be fully characterized" [3].
A model that flips on a single character
The most vivid demonstration involves code. The team asked DeepSeek-V4-Flash-Base to finish a line from its own FP8 kernel code, after the text "T.Cast(FP". FP8 is a compact eight-bit number format, and a kernel is a small piece of code that runs one specific calculation on the chip. The researchers then lengthened a docstring before the line one token at a time. The model's top choice flipped between the correct "8" and the wrong "32" every two tokens, repeating every four. That is the stride exactly [2][4]. Sixty of the 64 rankings followed the period-four pattern [4].
DeepSeek-V3.1-Base, which has no chunked compression, preferred "32" on only four of the same 64 inputs, each time by a probability margin below 0.07 [2][4].
A controlled test points at the compression itself
DeepSeek-V4 differs from older models in many ways, so the authors built a cleaner test. They pretrained 0.6-billion-parameter models from scratch on 100 billion tokens, using the Qwen3-0.6B architecture with chunked compression as the only substantial change from full-attention baselines [2][4].
The period followed the stride every time. Strides of four, six, eight and 12 produced periods of about four, six, eight and 12. Full-attention baselines stayed within 6.1 points across positions, while compressed models reached gaps of up to 78 points. That held even without rotary position embeddings, a common way of encoding where a token sits, and with plain averaging in place of learned gates, the adjustable weights that decide how much each token in a window counts toward its summary [2][4].
This is strong evidence about the technique, though it comes from small models the authors trained themselves, not from DeepSeek's production systems.
Averages hide the damage. One compressed model, with a window and a stride of eight tokens each, scored a mean of 59.4 percent against 61.1 percent for full attention. Its worst-position accuracy was 9.9 percent, against 58.4 percent for full attention [2].
Why the weak spots persist
Inside the models, different attention heads turn out to specialise in different phases. Heads are the parallel parts of each layer that decide which earlier tokens the model looks at. Knocking out a specific head cut accuracy at only two or three adjacent phases and left the rest alone [2][4].
The authors also shifted the learned gate settings by one slot in a cycle. The weak spots moved by exactly one position, matching the prediction almost perfectly, which the authors take as evidence that the gates set the pattern. The phase at which a compression window ends stayed weak under every shift [2].
Their theory points the same way. In an idealised model, the authors prove that concentrating each head on a single slot is optimal and that gradient flow, the continuous version of training by gradient descent, settles each head on a fixed slot. Nothing in the training dynamics pushes different heads toward different slots, so some phases can go uncovered [2][4]. If that holds in practice, the flaw is a product of how these models learn and not a tuning slip that a patch would clear.
What the paper asks for
The authors' recommendation is simple. Models with chunked compression should report long-context accuracy broken down by phase, not as one average score [2][4]. A model can look healthy on a leaderboard while failing systematically on facts that land at the wrong offset, as the 59.4 percent mean beside the 9.9 percent floor shows.
The evidence has limits. The tests measure retrieval of planted facts, and the paper does not show how much this costs in summarisation or in agents that resend long conversation histories. The effect has not been checked at larger scale. The work is a pre-print and has not been through peer review.
The authors have also published an interactive blog post where readers can shift the phase and explore the results model by model [4].
No response from DeepSeek to the paper has been found so far. Whether it, or any other lab that compresses this way, begins publishing phase-by-phase numbers is the thing to watch.
Every edition in brief, three times a day, on our Telegram channel, on Bluesky and on Threads.
Spotted an error? Tell the editors
- [2609.36322] Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression arxiv.org
- Paper page - Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression huggingface.co
- Eight Changes to Transformer Attention, and What Each Costs airealist.ai
- Periodic Weak Spots ultimatejupiter.github.io




