← Search

Rithesh Kumar

7 accepted papers

2026

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

ICML 2026poster

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound effects. Compared to traditional audio processing tasks, audio sto…

Cited by 0SourcecodeScholar
2026

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

ICLR 2026poster

While generative Text-to-Speech (TTS) systems leverage vast "in-the-wild" data to achieve remarkable success, speech-to-speech processing tasks like enhancement face data limitations, which lead data-hungry generative approaches to distort speech content and speaker identity. To bridge this gap, we…

Cited by 0SourceScholar
2025

DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis

ICML 2025poster

Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are computationally intensive, and previous distillation attempts have shown consistent quality degradation. Moreover, ex…

2023

High-Fidelity Audio Compression with Improved RVQGAN

NeurIPS 2023spotlight

Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natural signals into lower dimensional discrete tokens. To that end, we introduce a h…

2022

Chunked Autoregressive GAN for Conditional Waveform Synthesis

ICLR 2022poster

Conditional waveform synthesis models learn a distribution of audio waveforms given conditioning such as text, mel-spectrograms, or MIDI. These systems employ deep generative models that model the waveform via either sequential (autoregressive) or parallel (non-autoregressive) sampling. Generative a…

2019

MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis

NeurIPS 2019poster

Previous works (Donahue et al., 2018a; Engel et al., 2019a) have found that generating coherent raw audio waveforms with GANs is challenging. In this paper, we show that it is possible to train GANs reliably to generate high quality coherent waveforms by introducing a set of architectural changes an…

2017

SampleRNN: An Unconditional End-to-End Neural Audio Generation Model

ICLR 2017poster

In this paper we propose a novel model for unconditional audio generation task that generates one audio sample at a time. We show that our model which profits from combining memory-less modules, namely autoregressive multilayer perceptron, and stateful recurrent neural networks in a hierarchical str…

Cited by 761SourcecodeScholar