VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy conditions remains a significant and underexplored challenge in the…