← Search

Shivam Mehta

6 accepted papers

2026

GELINA: UNIFIED SPEECH AND GESTURE SYNTHESIS VIA INTERLEAVED TOKEN PREDICTION

ICASSP 2026oral

Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-spee…

Cited by 0SourcePDFScholar
2025

Make Some Noise: Towards LLM audio reasoning and generation using sound tokens

ICASSP 2025accepted

Integrating audio comprehension and generation into large language models (LLMs) remains challenging due to the continuous nature of audio and the resulting high sampling rates. Here, we introduce a novel approach that combines Variational Quantization with Conditional Flow Matching to convert audio…

Cited by 0SourceScholar
2024

Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching

ICASSP 2024accepted

We introduce Matcha-TTS, a new encoder-decoder architecture for speedy TTS acoustic modelling, trained using optimal-transport conditional flow matching (OT-CFM). This yields an ODE-based decoder capable of high output quality in fewer synthesis steps than models trained using score matching. Carefu…

Cited by 252SourceScholar
2024

Unified Speech and Gesture Synthesis Using Flow Matching

ICASSP 2024accepted

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associated body gestures. This paper presents a novel, unified architecture for jointl…

Cited by 7SourceScholar
2023

Prosody-Controllable Spontaneous TTS with Neural HMMS

ICASSP 2023accepted

Spontaneous speech has many affective and pragmatic functions that are interesting and challenging to model in TTS. However, the presence of reduced articulation, fillers, repetitions, and other disfluencies in spontaneous speech make the text and acoustics less aligned than in read speech, which is…

Cited by 0SourceScholar
2022

Neural HMMS Are All You Need (For High-Quality Attention-Free TTS)

ICASSP 2022accepted

Neural sequence-to-sequence TTS has achieved significantly better output quality than statistical speech synthesis using HMMs. However, neural TTS is generally not probabilistic and uses non-monotonic attention. Attention failures increase training time and can make synthesis babble incoherently. Th…

Cited by 0SourceScholar