Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models
Published in arXiv, 2026
Authors: Xuning He, Zinan Sheng, Yongding Tao, Huanyu Liu, Ge Li, Xue Jiang, Yihong Dong.
Abstract: Diffusion language models (DLMs) iteratively refine a sequence, allowing earlier predictions to be revised as context evolves. This rollback capability distinguishes them from irreversible autoregressive generation, but makes inference costly. In this paper, we introduce Archer, a training-free KV caching method for rollback-capable DLMs. Archer asymmetrically keeps the mutable response synchronized with the current hypothesis while reusing prompt K/V within a bounded state neighborhood. Archer attains the best mean performance of 33.63% together with a 2.57x mean speedup on the main suite, improves Pass@1 by up to 3.05 points and reaches up to 2.95x speedup.
arXiv Link: https://arxiv.org/abs/2608.08086
Recommended citation: Xuning He, Zinan Sheng, Yongding Tao, Huanyu Liu, Ge Li, Xue Jiang, Yihong Dong. "Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models." arXiv preprint arXiv:2608.08086 (2026).
Download Paper
