← Search

Nanda Dulal Jana

3 accepted papers

2025

KANGAN-AVSS: Kolmogorov-Arnold Network Based Generative Adversarial Networks for Audio-Visual Speech Synthesis

ICASSP 2025accepted

Audio-visual speech synthesis (AVSS) is an emerging research topic in the paradigm of generative AI, aiming to generate realistic and synchronized audio-visual outputs for a target speaker based on input audio from any source speaker, combining Voice Conversion (VC) and Audio-Visual Synthesis (AVS).…

Cited by 0SourceScholar
2025

SwinGAN-AVSS: Audio-Visual Speech Synthesis Leveraging Swin Transformer-Enhanced Generative Adversarial Networks

ICASSP 2025accepted

Audio-visual speech synthesis (AVSS) is a emerging field of study that involves generating synchronized and realistic video of a target speaker based on converted audio inputs of a source speaker. The AVSS method includes two sequential components: voice conversion (VC) to transform the source speak…

Cited by 0SourceScholar
2023

Voice Conversion Using Feature Specific Loss Function Based Self-Attentive Generative Adversarial Network

ICASSP 2023accepted

Voice conversion (VC) is the process of converting the vocal texture of a source speaker similar to that of a target speaker without altering the content of the source speaker’s speech. With the ongoing developments of deep generative models, generative adversarial networks (GANs) appeared as a bett…

Cited by 0SourceScholar