2025
Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
ICASSP 2025accepted
Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize reverberant speech for the spoken content. Previous works focus on the RGB modality for global environmental modeling, overlooking the potential of multi-source spatial knowledge like depth, speaker positio…