Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
Abstract
Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Training the generator on its own rollouts exposes it to these imperfect histories. However, existing video-level distribution matching distillation (DMD) scores the whole rollout jointly. Because a chunk is evaluated together with its past and future, its correction can favor matching artifacts in the surrounding context merely to preserve temporal consistency. To provide a clearer visual-quality signal, we introduce Rollout-Marginal Distillation (RMD). RMD retains the generated history for AR prediction but scores each chunk independently against a chunk teacher, ensuring its quality correction is not compromised by an imperfect temporal context. To compensate for the lack of temporal context in independent chunk scoring, RMD subsequently applies video-level DMD to restore temporal coherence. Extensive experiments demonstrate that RMD maintains high visual quality far beyond its training horizon and outperforms video-level DMD baselines. Code and video results are available at https://cjeen.github.io/RMD
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation (2026)
- Enhancing Autoregressive Video Generation via Representation Adversarial Distillation (2026)
- Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation (2026)
- Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout (2026)
- Uncertainty DMD: Restoring Diversity in Few-Step Autoregressive Video Distillation (2026)
- Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation (2026)
- ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper