ICASSP 2025accepted0 citations

SwinGAN-AVSS: Audio-Visual Speech Synthesis Leveraging Swin Transformer-Enhanced Generative Adversarial Networks

Subhayu Ghosh, Swapnil Saha, Nanda Dulal Jana

Abstract

Audio-visual speech synthesis (AVSS) is a emerging field of study that involves generating synchronized and realistic video of a target speaker based on converted audio inputs of a source speaker. The AVSS method includes two sequential components: voice conversion (VC) to transform the source speaker’s voice to the target speaker’s voice, and audio-visual synthesis (AVS) to generate synchronized video of the target speaker from the output of the VC model. This paper presents an AVSS approach using Swin Transformer-based generative adversarial network (GAN) framework. The Swin Transformer is incorporated into the discriminator of both the VC and AVS models. Its hierarchical design and self-attention mechanisms significantly enhance the temporal and spatial coherence of the synthesized outputs, thereby improving the quality and synchronization of both audio and visual components. Moreover, a feature matching loss in the VC model and a temporal coherence loss in the AVS model is also incorporated to enhance the quality of synthesized audio and video outputs. Experimental results demonstrate that the proposed approach significantly outperforms existing techniques in terms of audio quality and visual synchronization, as validated by objective metrics and subjective evaluations. This work advances AVSS, offering improved performance for applications in virtual avatars, dubbing, and human-computer interaction.

BibTeX
@inproceedings{icassp2025_swinganavssaudio,
  title = {SwinGAN-AVSS: Audio-Visual Speech Synthesis Leveraging Swin Transformer-Enhanced Generative Adversarial Networks},
  author = {Subhayu Ghosh and Swapnil Saha and Nanda Dulal Jana},
  booktitle = {ICASSP 2025},
  year = {2025}
}