← Search

Heng Lu

14 accepted papers

2026

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

ICLR 2026poster

Neural audio codecs are foundational to speech language models. It is expected to have a low frame rate and decoupled semantic and acoustic information. A lower frame rate codec can reduce the computational cost of speech language models by shortening the sequence length. Recent studies have develop…

Cited by 0SourcecodeScholar
2025

FlashAudio: Rectified Flow for Fast and High-Fidelity Text-to-Audio Generation

ACL 2025long

Recent advancements in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. While recent methods utilizing consistency-based distillation aim to achieve few-step…

2025

UniSpeaker: A Unified Approach for Multimodality-driven Speaker Generation

EMNLP 2025

While recent advances in reference-based speaker cloning have significantly improved the authenticity of synthetic speech, speaker generation driven by multimodal cues such as visual appearance, textual descriptions, and other biometric signals remains in its early stages. To pioneer truly multimoda

2024

Diacorrect: Error Correction Back-End for Speaker Diarization

ICASSP 2024accepted

In this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way. This method is inspired by error correction techniques in automatic speech recognition. Our model consists of two parallel convolutional encoders and…

Cited by 0SourceScholar
2024

GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition

ICASSP 2024accepted

Contrastive cross-modality pretraining has recently exhibited impressive success in diverse fields, whereas there is limited research on their merits in speech emotion recognition (SER). In this paper, we propose GEmo-CLAP, a kind of gender-attribute-enhanced contrastive language-audio pretraining (…

Cited by 0SourceScholar
2024

Promptvc: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

ICASSP 2024accepted

Stylistic voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or reference speech to control the conversion process, which leads to limitations…

Cited by 0SourceScholar
2023

Hybridformer: Improving Squeezeformer with Hybrid Attention and NSR Mechanism

ICASSP 2023accepted

SqueezeFormer has recently shown impressive performance in automatic speech recognition (ASR). However, its inference speed suffers from the quadratic complexity of softmax-attention (SA). In addition, limited by the large convolution kernel size, the local modeling ability of SqueezeFormer is insuf…

Cited by 0SourceScholar
2023

Symbolization, Prompt, and Classification: A Framework for Implicit Speaker Identification in Novels

EMNLP 2023long findings

Speaker identification in novel dialogues can be widely applied to various downstream tasks, such as producing multi-speaker audiobooks and converting novels into scripts. However, existing state-of-the-art methods are limited to handling explicit narrative patterns like "Tom said, '...'", unable to…

Cited by 0SourceScholar
2022

Improving Cross-Lingual Speech Synthesis with Triplet Training Scheme

ICASSP 2022accepted

Recent advances in cross-lingual text-to-speech (TTS) made it possible to synthesize speech in a language foreign to a monolingual speaker. However, there is still a large gap between the pronunciation of generated cross-lingual speech and that of native speakers in terms of naturalness and intellig…

Cited by 0SourceScholar
2022

The USTC-Ximalaya System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription (M2met) Challenge

ICASSP 2022accepted

We propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-Channel Multi-Party Meeting Transcription (M2MeT) challenge. These techniques are designed to handle multi-speaker conve…

Cited by 0SourceScholar
2020

Pitchnet: Unsupervised Singing Voice Conversion with Pitch Adversarial Network

ICASSP 2020accepted

Singing voice conversion is to convert a singer's voice to another one's voice without changing singing content. Recent work shows that unsupervised singing voice conversion can be achieved with an autoencoder-based approach [1]. However, the converted singing voice can be easily out of key, showing…

Cited by 0SourceScholar
2019

Enhancing Hybrid Self-attention Structure with Relative-position-aware Bias for Speech Synthesis

ICASSP 2019accepted

Compared with the conventional "front-end"-"back-end"- "vocoder" structure, based on the attention mechanism, end-to-end speech synthesis systems directly train and synthesize from text sequence to the acoustic feature sequence as a whole. Recently, a more calculation efficient end-to-end architectu…

Cited by 0SourceScholar
2018

Deep Feed-Forward Sequential Memory Networks for Speech Synthesis

ICASSP 2018accepted

The Bidirectional LSTM (BLSTM) RNN based speech synthesis system is among the best parametric Text-to-Speech (TTS) systems in terms of the naturalness of generated speech, especially the naturalness in prosody. However, the model complexity and inference cost of BLSTM prevents its usage in many runt…

Cited by 0SourceScholar