← Search

Vinay S. Raghavan

2 accepted papers

2025

Decoding the Unintelligible: Neural Speech Tracking in Low Signal-to-Noise Ratios

ICASSP 2025accepted

Understanding speech in noisy environments is challenging for both human listeners and speech technologies, with significant implications for hearing aid design and communication systems. Auditory attention decoding (AAD) aims to decode the attended talker from neural signals to enhance their speech…

Cited by 0SourceScholar
2023

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

NeurIPS 2023poster

In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis. StyleTTS 2 differs from its predecessor by modeling styles as a latent random variable through dif…