← Search

Jibin Song

1 accepted papers

2026

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

ICLR 2026poster

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video motion, making it a promising condition for temporally controlled video generatio…

Cited by 0SourceScholar