ICASSP 2025accepted0 citations

KANGAN-AVSS: Kolmogorov-Arnold Network Based Generative Adversarial Networks for Audio-Visual Speech Synthesis

Subhayu Ghosh, Swapnil Saha, Nanda Dulal Jana

Abstract

Audio-visual speech synthesis (AVSS) is an emerging research topic in the paradigm of generative AI, aiming to generate realistic and synchronized audio-visual outputs for a target speaker based on input audio from any source speaker, combining Voice Conversion (VC) and Audio-Visual Synthesis (AVS). The VC component transforms the audio characteristics of the source speaker to match the target speaker, while the AVS component generates a video that is synchronized with the converted audio. This paper introduces a novel approach to AVSS by incorporating the Kolmogorov Arnold Network (KAN) into the discriminator of Generative Adversarial Networks (GANs) for both VC and AVS tasks. By integrating KAN into the GAN framework, the proposed approach leverages KAN’s ability to model high-dimensional, non-linear relationships with fewer parameters, enhancing the discriminator’s efficiency and accuracy. This results in more natural synthesis of audio-visual outputs with reduced computational complexity and faster training time. Experimental evaluations on benchmark datasets demonstrate that our KAN-based GAN substantially outperforms existing methods in the quality and naturalness of generated samples. This work advances the state of AVSS by improving the realism of synthesized speech and video, offering a robust and efficient solution for creating synchronized audio-visual output.

BibTeX
@inproceedings{icassp2025_kanganavsskolmog,
  title = {KANGAN-AVSS: Kolmogorov-Arnold Network Based Generative Adversarial Networks for Audio-Visual Speech Synthesis},
  author = {Subhayu Ghosh and Swapnil Saha and Nanda Dulal Jana},
  booktitle = {ICASSP 2025},
  year = {2025}
}