MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation
Diffusion-based motion generation has advanced rapidly, but current methods still struggle with long-horizon consistency, style control, and multi-condition guidance. A major reason is the fused-conditioning design, where semantic, stylistic, and temporal signals share a single pathway, causing inte