EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This limitation stems from the intrinsic constraints of pixel-only training objectives, which inherently bias models toward app…