thu-ml/Causal-Forcing

[ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++

What it solves

Causal Forcing 解決了建立高品質、即時互動影片生成模型的挑戰。它特別針對自回歸 (AR) 擴散模型中的曝光偏差與運動動態退化問題,讓推論速度提升至 1‑4 步,同時不犧牲視覺品質或運動流暢度。

How it works

此專案實作了三階段訓練管線,將複雜的擴散模型蒸餾成快速的自回歸生成器:

  1. AR Diffusion Training:建立基礎的自回歸擴散模型。
  2. Causal Initialization:使用 Causal ODE(需要精選的 ODE 配對資料)或 Causal Consistency Distillation(Causal Forcing++,利用真實資料消除 ODE 資料整理需求)來產生理論上正確的初始化。
  3. Asymmetric DMD:套用 Distribution Matching Distillation 進一步降低生成所需的步數。

它同時支援 chunk-wiseframe-wise 模型,後者可透過將第一個潛在影格設為條件影像,原生支援 Image‑to‑Video (I2V) 生成。

Who it’s for

此框架適合從事影片生成的 AI 研究者與開發者,特別是想實作即時互動影片工具或長篇影片生成(如 Rolling Forcing 擴充)的使用者。

Highlights

  • Ultra-Low Latency:提供 1 步與 2 步的 frame‑wise 模型,實現極速生成。
  • High Fidelity:在視覺品質與運動動態上均優於先前方法如 Self Forcing。
  • Versatile Modalities:支援 Text‑to‑Video (T2V) 與 Image‑to‑Video (I2V)。
  • Long Video Support:相容 Rolling Forcing 等技術,可生成分鐘級別的影片。