← Search

Zeyu Jin

27 accepted papers

2026

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

ICML 2026poster

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound effects. Compared to traditional audio processing tasks, audio sto…

Cited by 0SourcecodeScholar
2026

DITSE: HIGH-FIDELITY GENERATIVE SPEECH ENHANCEMENT VIA LATENT DIFFUSION TRANSFORMERS

ICASSP 2026poster

Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approaches for speech enhancement have shown promising results, they still face two ma…

Cited by 0SourcePDFScholar
2026

From Natural Alignment to Conditional Controllability in Multimodal Dialogue

ICLR 2026poster

The recent advancement of Artificial Intelligence Generated Content (AIGC) has led to significant strides in modeling human interaction, particularly in the context of multi-modal dialogue. While current methods impressively generates realistic dialogue in speech and vision modalities, challenges r…

Cited by 0SourcecodeScholar
2026

GENCHO: ROOM IMPULSE RESPONSE GENERATION FROM REVERBERANT SPEECH AND TEXT VIA DIFFUSION TRANSFORMERS

ICASSP 2026oral

Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability and degraded performance under unseen conditions. Moreover, emerging generative audio applications call for more flexible…

Cited by 0SourcePDFScholar
2026

PROMPTSEP: GENERATIVE AUDIO SEPARATION VIA MULTIMODAL PROMPTING

ICASSP 2026oral

Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separa…

Cited by 0SourcePDFScholar
2026

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

ICLR 2026poster

While generative Text-to-Speech (TTS) systems leverage vast "in-the-wild" data to achieve remarkable success, speech-to-speech processing tasks like enhancement face data limitations, which lead data-hungry generative approaches to distort speech content and speaker identity. To bridge this gap, we…

Cited by 0SourceScholar
2025

Code Drift: Towards Idempotent Neural Audio Codecs

ICASSP 2025accepted

Neural codecs have demonstrated strong performance in high-fidelity compression of audio signals at low bitrates. The token-based representations produced by these codecs have proven particularly useful for generative modeling. While much research has focused on improvements in compression ratio and…

Cited by 0SourceScholar
2025

DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis

ICML 2025poster

Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are computationally intensive, and previous distillation attempts have shown consistent quality degradation. Moreover, ex…

2025

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference

ICASSP 2025accepted

We propose a unified framework for Singing Voice Synthesis (SVS) and Conversion (SVC), addressing the limitations of existing approaches in cross-domain SVS/SVC, poor output musicality, and scarcity of singing data. Our framework enables control over multiple aspects, including language content base…

Cited by 0SourceScholar
2025

Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs

ICLR 2025poster

Large Vision-Language Models (LVLMs) often produce responses that misalign with factual information, a phenomenon known as hallucinations. While hallucinations are well-studied, the exact causes behind them remain underexplored. In this paper, we first investigate the root causes of hallucinations i…

2024

A Closer Look at the Limitations of Instruction Tuning

ICML 2024poster

Instruction Tuning (IT), the process of training large language models (LLMs) using instruction-response pairs, has emerged as the predominant method for transforming base pre-trained LLMs into open-domain conversational agents. While IT has achieved notable success and widespread adoption, its limi…

Cited by 18SourcePDFScholar
2024

GR0: Self-Supervised Global Representation Learning for Zero-Shot Voice Conversion

ICASSP 2024accepted

Research in generative self-supervised learning (SSL) has largely focused on local embeddings for tokenized sequences. We introduce a generative SSL framework that learns a global representation that is disentangled from local embeddings. We apply this technique to jointly learn a global speaker emb…

Cited by 0SourceScholar
2024

MDX-GAN: Enhancing Perceptual Quality in Multi-Class Source Separation Via Adversarial Training

ICASSP 2024accepted

Audio source separation aims to extract individual sound sources from an audio mixture. Recent studies on source separation focus primarily on minimizing signal-level distance, typically measured by source-to-distortion ratio (SDR). However, scant attention has been given to the perceptual quality o…

Cited by 0SourceScholar
2022

Controllable Speech Representation Learning Via Voice Conversion and AIC Loss

ICASSP 2022accepted

Speech representation learning transforms speech into features that are suitable for downstream tasks, e.g. speech recognition, phoneme classification, or speaker identification. For such recognition tasks, a representation can be lossy (non-invertible), which is typical of BERT-like self-supervised…

Cited by 0SourceScholar
2021

CDPAM: Contrastive Learning for Perceptual Audio Similarity

ICASSP 2021accepted

Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function. The DPAM approach of Manocha et al. [1] learns a full-reference metric trained directly on human judgments, and thus correlates well with human perception. However, it re…

Cited by 0SourceScholar
2021

Context-Aware Prosody Correction for Text-Based Speech Editing

ICASSP 2021accepted

Text-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major drawback of current systems, however, is that edited recordings often sound unnatural because of prosody mismatches around e…

Cited by 0SourceScholar
2020

Disentangled Multidimensional Metric Learning for Music Similarity

ICASSP 2020accepted

Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessary to define a similarity metric to compare one recording to another. Music simi…

Cited by 0SourceScholar
2020

F0-Consistent Many-To-Many Non-Parallel Voice Conversion Via Conditional Autoencoder

ICASSP 2020accepted

Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Many style-transfer-inspired methods such as generative adversarial networks (GANs) and variational autoencoders (VAEs) have been proposed. Recently, AutoVC, a conditional autoencoders (CAEs) ba…

Cited by 0SourceScholar
2016

Cute: A concatenative method for voice conversion using exemplar-based unit selection

ICASSP 2016accepted

State-of-the art voice conversion methods re-synthesize voice from spectral representations such as MFCCs and STRAIGHT, thereby introducing muffled artifacts. We propose a method that circumvents this concern using concatenative synthesis coupled with exemplar-based unit selection. Given parallel sp…

Cited by 0SourceScholar