MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural g