AAAI 2026technical0 citations

MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer

Penghui Liu, Jiangshan Wang, Yutong Shen, Shanhui Mo, Chenyang Qi, Jack Ma

Abstract

Multi-object video motion transfer poses significant challenges for Diffusion Transformer (DiT) architectures due to inherent motion entanglement and lack of object-level control. We present MultiMotion, a novel unified framework that overcomes these limitations. Our core innovation is Mask-aware Attention Motion Flow (AMF), which utilizes SAM 2 masks to explicitly disentangle and control motion features for multiple objects within the DiT pipeline. Furthermore, we introduce RectPC, a high-order predictor-corrector solver for efficient and accurate sampling, particularly beneficial for multi-entity generation. To facilitate rigorous evaluation, we construct the first benchmark dataset specifically for DiT-based multi-object motion transfer. MultiMotion demonstrably achieves precise, semantically aligned, and temporally coherent motion transfer for multiple distinct objects, maintaining DiT

BibTeX
@inproceedings{aaai2026_multimotionmulti,
  title = {MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer},
  author = {Penghui Liu and Jiangshan Wang and Yutong Shen and Shanhui Mo and Chenyang Qi and Jack Ma},
  booktitle = {AAAI 2026},
  year = {2026}
}
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer · AAAI 2026