Freezing the decoder is a better way to handle SVD-based KV-cache compression
Using a shared learning rate to compare different parameter-efficient fine-tuning recipes is often a mistake. When comparing methods with vastly different numbers of trainable parameters, a single rate can suppress the performance of larger sets while inflating the variance, making the smallest parameter set look like a winner when it isn't. This is clearly seen in post-hoc SVD-based KV-cache compression, where pretrained models are converted to low-rank caches by splitting key/value weights into a down-projection encoder and an up-projection decoder. While a short "healing" fine-tune is used to recover accuracy lost during truncation, the common practice of using one learning rate across all tests creates a false advantage for encoder-only healing.
How does encoder-only healing actually perform?
In the specific context of SVD-based compression, the weights are factorized into two parts: the encoder (down-projection) and the decoder (up-projection). Initial observations suggested that freezing the decoder and only training the encoder was superior to training the decoder or both. However, this advantage is an artifact of the shared learning rate. When each configuration is given its own tuned learning rate, encoder-only healing actually reaches parity with the other methods.
The real value here is efficiency. By freezing the decoder, the process achieves a measured saving of 3x fewer trainable parameters and 3x less optimizer-state memory. This makes it a viable, low-memory drop-in recipe for anyone retrofitting low-rank KV-cache compression during training.
Why does the shared learning rate fail?
The failure happens because the "arms" (the different training configurations) have different trainable-parameter counts. A learning rate that works for a small set of parameters might be suboptimal for a larger set. When a shared rate is forced upon both, the larger set's mean performance is depressed and its variance is inflated. This manufactures a statistical advantage for the smaller arm that does not exist in reality.
To prove this, the research verified the parity across different setups:
- Using Qwen2.5-VL-3B-Instruct (a vision-language model) at a specific compression ratio.
- Running three seeds per configuration to ensure the results weren't fluke.
- Replicating the findings on a text-only testbed using two different backbones.
What should be done when applying this to a model?
If you are attempting to recover accuracy after SVD-based truncation of KV-caches, the logical step is to prioritize the encoder for healing. Since the performance is parity-equivalent to training both components—but requires significantly fewer resources—there is no reason to waste memory on the decoder.
The critical failure point in most implementation pipelines is the evaluation phase. If you are comparing different fine-tuning subsets (like LoRA ranks or specific layer freezes), you cannot use a single learning rate for the benchmark. To avoid the "shared-learning-rate pitfall," each subset must be tuned independently. If you see a massive performance jump in a version with very few trainable parameters, check if the learning rate was optimized specifically for that parameter count. If it wasn't, and you're using a shared rate, the "win" is likely an illusion.
For those implementing this on the Qwen2.5-VL-3B-Instruct or similar architectures, the workflow should be:
- Perform SVD factorization on the key/value weights.
- Set the decoder (up-projection) to frozen.
- Apply a tuned learning rate specifically for the encoder (down-projection) during the healing phase.
- Verify that accuracy has returned to acceptable levels without inflating the optimizer-state memory.
The full details of this method can be found at https://arxiv.org/abs/2610.10552.
Does freezing the decoder mean the SVD factors are computed once and kept static during inference, or do they adapt to the input context?