HomeArticle

ByteDance has found the reason why DeepSeek's performance fluctuates between strong and weak.

量子位2026-10-09 15:27
Whether the answer is correct depends on the token's position.

The bipolar performance quirk of DeepSeek that alternates between exceptional and dysfunctional outputs has been spotted by ByteDance's Seed team.

For the exact same problem with no content modified, if you just add a few irrelevant characters at the beginning, the model suddenly fails to solve it???

And this is not an occasional random glitch.

Seed researchers found that the performance of DeepSeek-V4 surprisingly changes periodically every 4 Tokens depending on the input position of the information.

What does this mean? Whether the model can remember a piece of information actually depends on where the information is placed in the input??

Add 2 more Tokens at the front, the answer may change from wrong to correct; add another 2, it will turn back to the wrong one again.

Alright alright, even model answering performance starts to depend on "positioning" now.

In the 128K long-context retrieval test, researchers found that when only the position of the same piece of information is changed, the retrieval accuracy of the DeepSeek-V4 series models can differ by up to 40.2 percentage points.

The team further found that this issue is related to a long-context optimization technique adopted by DeepSeek-V4 --

Chunked KV Cache Compression.

This technique is originally designed to make the model process long texts more memory-efficient and faster.

As a result, after compression, the model ends up switching between being surprisingly capable and completely useless (doge).

Add 2 Extra Tokens, DeepSeek Suddenly Gets the Answer Right

ByteDance Seed researchers first conducted an experiment using DeepSeek's own code.

The test subject was DeepSeek-V4-Flash-Base.

They intercepted a section of FP8 quantization function from DeepSeek-V4's official inference code, and asked the model to complete the last Token.

The correct answer to the task should be 8, since this section of code is supposed to implement FP8-related type conversion.

But sometimes the model insists that the correct completion should be 32.

To figure out what exactly went wrong, researchers added a section of purely decorative docstring in front of the code, filled with repeated equal signs.

After that, they started adjusting the number of equal signs.

The code itself was not modified, the position to be completed was not changed, and the correct answer certainly stayed the same... The only change was those few Tokens with no practical meaning added at the front.

As a result, DeepSeek's answer started to flip back and forth repeatedly.

When the padding length falls in certain positions, the model tends to output the wrong answer 32.

Move it 1 or 2 Tokens forward, it will tend to output the correct answer 8 again.

Keep moving it, the wrong answer comes back --

The entire process repeats with a cycle of 4 Tokens.

More specifically, among the 16 padding lengths tested in the paper, when the remainder of the length divided by 4 is 0 or 1, the model tends to output the wrong answer 32; when the remainder is 2 or 3, it tends to output the correct answer 8.

Researchers also calculated the probability the model assigns to the two candidate answers.

At one set of positions, the average probability of the wrong answer 32 reaches 71.3%, while the correct answer 8 only accounts for 26.4%.

At another set of positions, the situation is completely reversed:

The average probability of the correct answer 8 rises to 91.5%, and the wrong answer 32 only takes 7.2%.

That is to say, the length change of the first few irrelevant characters is enough to make the model produce completely different judgments on the exact same question.

This is rather hard to comment on.

When programmers debug code, they usually check the logic, variables and dependencies first.

Now it seems that they may also need to check by the way if they have typed two extra equal signs at the front.

However, a single code completion case is not enough to show how widespread this problem is.

So the Seed team continued to expand the test scope.

They turned to a very classic task in the long-context capability test of large models, Needle-in-a-Haystack.

Researchers constructed a context as long as 128K Tokens, which contains about 16,000 key-value pairs.

For example, K1 corresponds to V1, K2 corresponds to V2...

Then they asked the model to find the Value corresponding to the specified Key from the context.

During the test, the key-value relationship remains unchanged, the question remains unchanged, and the total length of the context also stays consistent.

Researchers focused on adjusting the position of the target information relative to the boundary of the compression window.

The results show that the accuracy curve of the DeepSeek-V4 series presents very obvious periodic fluctuations.

The maximum accuracy gap of DeepSeek-V4-Flash-Base between different positions reaches 40.2 percentage points

DeepSeek-V4-Pro-Base also reaches 34.8 percentage points.

After post-training, the situation has improved to some extent.

The gap of DeepSeek-V4-Flash-0731 is reduced to 19.1 percentage points, and that of DeepSeek-V4-Pro-0813 is reduced to 14.8 percentage points.

For the newer DeepSeek-V4.1-Flash-0910, the gap is further reduced to 6.1 percentage points.

But the periodic difference still exists.

Where does this difference come from?

Looking closer, the fluctuation cycle of DeepSeek-V4 is 4 Tokens, while that of DeepSeek-V4.1 becomes 2 Tokens.

Researchers found that this exactly corresponds to the respective KV Cache compression step size adopted by the two generations of models.

Wow, even the fluctuation cycle of the answering performance matches the underlying compression configuration.

Is the Root Cause in KV Cache Compression?

Now let's talk about KV Cache compression.

When processing long contexts, large models need to store a large amount of Key and Value information corresponding to historical Tokens for subsequent attention calculation.

The longer the context, the more memory and computing overhead this part of cache takes up.

Especially for long tasks with hundreds of thousands or millions of Tokens, KV Cache can easily become the bottleneck of inference efficiency.

Therefore, DeepSeek-V4 adopts Chunked KV Cache Compression.

The idea is to divide consecutive Tokens into separate windows, and then compress the information in the window into fewer cache entries.

In this way, the model does not need to retain the same scale of cache for each historical Token.

It not only saves memory, but also reduces the cost of long-context attention calculation.

But the Seed team found that the reason why DeepSeek shows the alternating brilliant and dysfunctional performance may hide in this chunking process.

Assume that every 4 Tokens form a compression step size.

Then, when the same piece of information appears at the 1st, 2nd, 3rd or 4th position in the window, the conditions under which it is compressed may be different.

The paper refers to this position relative to the boundary of the compression window as Phase.

Researchers found that the model has systematic differences in retrieval capabilities for information in different phases.

They named this phenomenon Phase Sensitivity.

For example, give the same set of materials to the model:

In the first typesetting, the key number just falls in the position that the model is easy to retain;

In the second typesetting, only a few more words are added at the front, which changes the position of the key number relative to the compression window.

The materials are still the same, but whether the model can find the number later may show obvious difference.

Moreover, this problem cannot be simply attributed to "the information is just cut between two windows".

Researchers found that even if both the Key and Value fall in the same compression window, the retrieval accuracy at different positions can still vary greatly.

This shows that the problem also involves how the model writes information into the compressed cache and how it reads data from the cache later.

To further confirm this conclusion, the Seed team simply trained a batch of models from scratch.

Based on the Qwen3-0.6B architecture, they constructed multiple KV Cache compression schemes, and set the full-attention model without chunked compression as the control group.

They only focused on changing the compression mechanism to see if periodic fluctuations would appear accordingly.

As a result, all tested chunked compression models show periodic changes corresponding to the compression step size.

On the contrary, the full-attention baseline model does not show the same degree of periodicity.

Researchers also adjusted the window size and compression step size respectively, and found that the cycle mainly changes with the compression step size.

When the step size is 4, the performance fluctuates with a cycle of about 4 Tokens;

When the step size is 6, the cycle also becomes about 6 Tokens;

When the step size is 8, the same rule applies;

...

Even if RoPE positional encoding is not used, or the learnable compression weights are replaced with simple averaging, this phenomenon still exists.

In other words, the problem is not caused by a certain positional encoding or a certain special module alone.

The design of chunked compression itself may introduce periodic retrieval weaknesses.

The Seed team further intervened on the attention heads, and observed how the model's retrieval capability at each phase changes after removing different components.

The results show that different attention heads contribute differently to different phases.

Some heads are better at processing information in certain positions, while other heads play a greater role in other positions.

Researchers call this Phase Specialization.

That is to say, a division of labor seems to have formed inside the model: different attention components have developed preferences for different positions in the compression window.

This division of labor can help the model complete information retrieval, but it may also make some positions relatively weak links.

The paper also analyzed the training process through a simplified theoretical model, and found that gradient flow may push the compression module to form a stable positional preference.

This also explains why this periodicity is not simple random noise. It may be a natural outcome when the model learns how to compress information.

However, this kind of problem may not be directly observed from ordinary Benchmarks, because conventional evaluations usually aggregate a large number of test results into an average score.

Assuming that the model performs very well in some positions and lags significantly in other positions, the final calculated average score may still be good.

Therefore, the Seed team proposes that when evaluating models that adopt Chunked KV Cache Compression, we should not only look at the overall retrieval accuracy, but also place the same piece of information in different compression phases and test it respectively.

Even if there is such position-induced performance deviation, given the memory and computing cost of long-context inference, compression still has very practical value.

But after saving cache, the model may no longer treat information in different positions equally.

Of course, according to the results of this experiment, post-training and architecture iteration can indeed significantly narrow this gap.

Paper link: https://arxiv.org/pdf/2609.36322

This article is from the WeChat official account QuantumBit, author: Wen Le, published with authorization from 36Kr.