← Search

Jiaqi Su

15 accepted papers

2026

DITSE: HIGH-FIDELITY GENERATIVE SPEECH ENHANCEMENT VIA LATENT DIFFUSION TRANSFORMERS

ICASSP 2026poster

Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approaches for speech enhancement have shown promising results, they still face two ma…

Cited by 0SourcePDFScholar
2026

GENCHO: ROOM IMPULSE RESPONSE GENERATION FROM REVERBERANT SPEECH AND TEXT VIA DIFFUSION TRANSFORMERS

ICASSP 2026oral

Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability and degraded performance under unseen conditions. Moreover, emerging generative audio applications call for more flexible…

Cited by 0SourcePDFScholar
2026

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

ICLR 2026poster

While generative Text-to-Speech (TTS) systems leverage vast "in-the-wild" data to achieve remarkable success, speech-to-speech processing tasks like enhancement face data limitations, which lead data-hungry generative approaches to distort speech content and speaker identity. To bridge this gap, we…

Cited by 0SourceScholar
2025

Unraveling the Effects of Synthetic Data on End-to-End Autonomous Driving

ICCV 2025poster

End-to-end (E2E) autonomous driving (AD) models require diverse, high-quality data to perform well across various driving scenarios. However, collecting large-scale real-world data is expensive and time-consuming, making high-fidelity synthetic data essential for enhancing data diversity and model r…

2024

GR0: Self-Supervised Global Representation Learning for Zero-Shot Voice Conversion

ICASSP 2024accepted

Research in generative self-supervised learning (SSL) has largely focused on local embeddings for tokenized sequences. We introduce a generative SSL framework that learns a global representation that is disentangled from local embeddings. We apply this technique to jointly learn a global speaker emb…

Cited by 0SourceScholar
2024

MDX-GAN: Enhancing Perceptual Quality in Multi-Class Source Separation Via Adversarial Training

ICASSP 2024accepted

Audio source separation aims to extract individual sound sources from an audio mixture. Recent studies on source separation focus primarily on minimizing signal-level distance, typically measured by source-to-distortion ratio (SDR). However, scant attention has been given to the perceptual quality o…

Cited by 0SourceScholar
2024

TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling

COLING 2024main

As a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requires not only modeling the relationships between objects in the image but also mining the connections between adjacent im…

Cited by 1SourcePDFScholar
2022

Controllable Speech Representation Learning Via Voice Conversion and AIC Loss

ICASSP 2022accepted

Speech representation learning transforms speech into features that are suitable for downstream tasks, e.g. speech recognition, phoneme classification, or speaker identification. For such recognition tasks, a representation can be lossy (non-invertible), which is typical of BERT-like self-supervised…

Cited by 0SourceScholar