UniCache BAGEL-generated aurora over a red cabin beside a frozen fjord

UniCache

Task- and Type-Aware KV Cache Compression for Unified Multimodal Models

Generated by UniCache × Bagel

Wanqi Yang1,2,3Yuexiao Ma4Mei Xie5Xiawu Zheng6Shiwei Liu1,2,3

1 Max Planck Institute for Intelligent Systems · 2 ELLIS Institute Tübingen · 3 Tübingen AI Center · 4 Nanyang Technological University · 5 Independent Researcher · 6 Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Xiamen University

Task-Dependent Cache Composition

In MoT-based unified multimodal models, the information required by different tasks is represented by distinct KV cache types.

01

Image understanding

Image understanding uses instruction and source-ViT KVs to represent the instruction and the source image's semantic information, respectively.

02

Text-to-image generation

Text-to-image generation conditions on instruction KV.

03

Image editing

Image editing uses instruction, source-ViT, and source-VAE KVs together, with source-VAE KV providing source-image details.

Three tasks in a unified multimodal model activate different combinations of instruction, source-ViT, source-VAE, generated-text, and current-latent KV
Examples of deploying a Mixture-of-Transformers unified multimodal model for multi-task inference and the corresponding task-dependent KV cache.
Two image-editing comparisons showing input, full KV, eviction baselines, KIVI quantization, and UniCache outputs
The editing examples illustrate UniCache's advantages over uniformly applied compression policies at a target compression rate of 80%.

UniCache

UniCache identifies the KV cache types present in each task and separates compression candidates from protected segments. Through offline calibration of attention distributions, UniCache assigns suitable compression policies to different candidate types.

1

Task-Aware Cache Segmentation

UniCache partitions the attention context into segments based on the current task and semantic role.

2

Offline Policy Assignment

UniCache assigns a compression policy to each candidate task--type pair using attention sparsity measured through offline calibration.

3

Attention-Guided Dynamic Budget Allocation

UniCache estimates cache importance by aggregating attention probabilities over chunks of layers and inference steps, reducing scheduling frequency.

4

Type-Aware Parallel Compression

Once configured, each policy maintains its own runtime state and compresses only its assigned segments, allowing the segment compression to execute independently and in parallel.

UniCache framework: offline policy assignment, task-aware segmentation, attention-guided budget allocation, parallel compression, and KV reassembly
UniCache identifies the cache segments activated by each task, assigns suitable compression policies via offline calibration and allocates type-specific compression budgets from the corresponding attention statistics.

Empirical Results

UniCache effectively mitigates the failure modes encountered by prior KV cache compression methods and achieves balanced quality in all evaluated tasks through task-aware and type-aware KV cache compression.

BAGEL

MethodMME ↑GenEval ↑PIE Structure ↓PIE PSNR ↑
Full KV2373.210.7810.10118.805
Global H₂O2372.600.7770.12514.663
Global KIVI2362.310.7770.09818.576
UniCache2377.360.7800.10018.969

SenseNova-U1

MethodMME ↑MMMU ↑GenEval ↑PIE Structure ↓
Full KV2396.5047.670.8870.099
Global H₂O2290.2747.670.8750.142
Global KIVI2377.6547.000.8880.131
UniCache2383.7248.000.8920.100

Up to 5×logical KV compression in understanding and editing

Up to 1.78×throughput in the reported long-context, single-A100 physical-engine setting

Get Started

The repository includes a PyTorch reference implementation, a physical CUDA engine, and a full-KV baseline. Model weights are obtained separately.

Installation and backend details ↗
Image understanding · PyTorch backend
python infer.py --backend torch --task understanding \
  --model-path checkpoints/BAGEL-7B-MoT \
  --image /path/to/source.jpg \
  --prompt "Describe the image." \
  --output-dir outputs/understanding