← Search

Tan Dat Nguyen

6 accepted papers

2026

MAGE: A COARSE-TO-FINE SPEECH ENHANCER WITH MASKED GENERATIVE MODEL

ICASSP 2026poster

Speech enhancement remains challenging due to the trade-off between efficiency and perceptual quality. In this paper, we introduce MAGE, a Masked Audio Generative Enhancer that advances generative speech enhancement through a compact and robust design. Unlike prior masked generative models with rand…

Cited by 0SourcePDFScholar
2026

SPADE: STRUCTURED PRUNING AND ADAPTIVE DISTILLATION FOR EFFICIENT LLM-TTS

ICASSP 2026oral

The goal of this paper is to introduce SPADE, a framework for Structured Pruning and Adaptive Distillation for Efficient Large Language Model-based text-to-speech (LLM-TTS). Recent LLM-TTS systems achieve strong controllability and zero-shot generalization, but their large parameter counts and high…

Cited by 0SourcePDFScholar
2025

Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding

ICASSP 2025accepted

The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexible trade-offs between speed and quality during inference without requiring additional training. Our core idea is to pre…

Cited by 0SourceScholar
2025

AdaptVC: High Quality Voice Conversion with Adaptive Learning

ICASSP 2025accepted

The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to extract disentangled linguistic content from the source and voice style from the reference. While existing approaches lev…

Cited by 0SourceScholar
2025

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

ICASSP 2025accepted

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy conditions remains a significant and underexplored challenge in the…

Cited by 0SourceScholar
2024

Fregrad: Lightweight and Fast Frequency-Aware Diffusion Vocoder

ICASSP 2024accepted

The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps…

Cited by 0SourceScholar