← Search

Tianrui Wang

14 accepted papers

2026

EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis

ICASSP 2026poster

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language model (LLM)-based designs, rely on scaling fixed emotion embeddings or external…

Cited by 0SourcePDFScholar
2026

Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration

ICRA 2026poster

Functional grasping with dexterous robotic hands is a key capability for enabling tool use and complex manipulation, yet progress has been constrained by two persistent bottlenecks: the scarcity of large-scale datasets and the absence of integrated semantic and geometric reasoning in learned models.…

2026

InstructAudio: Unified speech and music generation with natural language instruction

ICASSP 2026poster

Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only limited text-level attribute control, and rarely support dialogue generation. TTM systems are constrained by input condi…

Cited by 0SourcePDFScholar
2025

A Chinese Expressive Long-dialogue Speech Dataset with Scripts

ICASSP 2025accepted

With the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervis…

Cited by 0SourceScholar
2025

Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding

ICASSP 2025accepted

Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. Adaptation on the pre-trained multi-lingual model has shown promising performance for CS-ASR. In this paper, we adapt Whi…

Cited by 0SourceScholar
2025

Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge

ICASSP 2025accepted

In this paper, we present the system developed by our team, CCATTS, for the LIMMITS’25 challenge, focusing on few-shot and zero-shot TTS. We adopt a two-stage TTS strategy. In track 1, we fine-tune the pre-trained ZMM-TTS model and successfully achieve multilingual low-latency TTS. In track 2, we pr…

Cited by 0SourceScholar
2025

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

NeurIPS 2025poster

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through ite…

Cited by 0SourcecodeScholar
2025

Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement

ICASSP 2025accepted

In recent speech enhancement (SE) research, transformer and its variants have emerged as the predominant methodologies. However, the quadratic complexity of the self-attention mechanism imposes certain limitations on practical deployment. Mamba, as a novel state-space model (SSM), has gained widespr…

Cited by 0SourceScholar
2025

Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module

ICASSP 2025accepted

The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coef…

Cited by 0SourceScholar
2025

Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement

ICASSP 2025accepted

Time-frequency (T-F) domain methods for monaural speech enhancement have benefited from the success of deep learning. Recently, focus has been put on designing two-stream network models to predict amplitude mask and phase separately, or, coupling the amplitude and phase into Cartesian coordinates an…

Cited by 0SourceScholar
2025

Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis

NeurIPS 2025spotlight

While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modelin…

Cited by 0SourceScholar
2023

An Adapter Based Multi-Label Pre-Training for Speech Separation and Enhancement

ICASSP 2023accepted

In recent years, self-supervised learning (SSL) has achieved tremendous success in various speech tasks due to its power to extract representations from massive unlabeled data. However, compared with tasks such as speech recognition (ASR), the improvements from SSL representation in speech separatio…

Cited by 0SourceScholar
2022

HGCN: Harmonic Gated Compensation Network for Speech Enhancement

ICASSP 2022accepted

Mask processing in the time-frequency (T-F) domain through the neural network has been one of the mainstreams for single-channel speech enhancement. However, it is hard for most models to handle the situation when harmonics are partially masked by noise. To tackle this challenge, we propose a harmon…

Cited by 0SourceScholar
2022

Harmonic Gated Compensation Network Plus for ICASSP 2022 DNS Challenge

ICASSP 2022accepted

The harmonic structure of speech is resistant to noise, but the harmonics may still be partially masked by noise. Therefore, we previously proposed a harmonic gated compensation network (HGCN) to predict the full harmonic locations based on the unmasked harmonics and process the result of a coarse e…

Cited by 0SourceScholar