← Search

Shuwei He

2 accepted papers

2025

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

AAAI 2025technical

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many attempts have been made to extract global spatial visual informat…