← Reading list
June 26, 2026

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models

Authors
Kaiwen Zheng, Guande He, Min Zhao, Jintao Zhang, et al.
Venue
arXiv 2606.25473
Link
Open →
Tags
world-modelsvideo-diffusiondistillationefficiency

Core idea

Extends rCM to the autoregressive video setting: TF-based continuous-time CMs + SF-based DMD, via a custom masked FlashAttention-2 JVP kernel for efficient causal training. ~10× faster convergence; 2-step generation on a 1.3B model; applied to streaming video + action-conditioned world models (Cosmos). Hits my world-models + inference-efficiency interests at once.