Papers
arxiv:2607.17715

C^2KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference

Published on Jul 20
Authors:
,
,
,
,
,
,
,
,
,

Abstract

C²KV enables efficient non-prefix KV cache reuse and compression for long-context LLM inference by learning position-agnostic compressed representations and co-training extraction with concatenation.

Long-context inference is central to modern large language model (LLM) applications such as retrieval-augmented generation and multi-document reasoning. To mitigate the growing inference cost, recent work has explored key-value (KV) cache reuse to reduce redundant prefill computation. However, existing reuse methods primarily focus on computation savings and overlook a critical bottleneck in long-context LLM serving: the cost of storing and accessing large KV caches. While KV compression appears to be a natural complement, naively combining compression with non-prefix KV reuse often leads to severe accuracy degradation. In this work, we propose C^2KV, a unified framework for non-prefix KV reuse that jointly optimizes KV extraction and inference-time concatenation. C^2KV learns a composable and compressed KV cache manifold that is explicitly designed to be position-agnostic. Our approach introduces a lightweight sidecar Extractor with learnable compression tokens and a structured attention flow, enabling modular KV representations that can be flexibly reused and concatenated without modifying the frozen base model. We further employ a compression-concatenation co-training strategy to align extraction-time representations with their downstream reuse behavior. Extensive experiments across multiple long-context benchmarks and model families demonstrate that C^2KV significantly reduces KV cache storage and transfer costs, achieving up to 17times inference speedup under long contexts, while preserving generation quality.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.17715
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.17715 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.17715 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.