← Search

Zhiyong Wu

124 accepted papers

2026

DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models

AAAI 2026technical

Extending pre-trained text Large Language Models (LLMs)’s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention in the speech research community. However, building a unified speech understanding and generation model still faces the

Cited by 0SourcePDFScholar
2026

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

ICML 2026poster

Audio-driven visual dubbing aims to synchronize a video's lip movements with new speech but is fundamentally challenged by the lack of ideal training data: paired videos differing only in lip motion. Existing methods circumvent this via mask-based inpainting. However, masking inevitably destroys spa…

Cited by 0SourceScholar
2026

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

ICLR 2026poster

Generative models for speech synthesis face a fundamental trade-off: discrete tokens ensure stability but sacrifice expressivity, while continuous signals retain acoustic richness but suffer from error accumulation due to task entanglement. This challenge has driven the field towards multi-stage pip…

Cited by 0SourcecodeScholar
2026

Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

AAAI 2026technical

Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, images, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of modality-complete data and the difficulty

Cited by 0SourcePDFScholar
2026

PROMO: Promptable Outfitting for Efficient High-Fidelity Virtual Try-On

CVPR 2026

Virtual Try-on (VTON) has become a core capability for online retail, where realistic try-on results provide reliable fit guidance, reduce returns, and benefit both consumers and merchants. Diffusion-based VTON methods achieve photorealistic synthesis, yet often rely on intricate architectures such

Cited by 0SourceScholar
2026

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

ICLR 2026poster

Large Language Models (LLMs) have extended their impact beyond Natural Language Processing, substantially fostering the development of interdisciplinary research. Recently, various LLM-based agents have been developed to assist scientific discovery progress across multiple aspects and domains. Among…

Cited by 0SourcecodeScholar
2026

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

ICML 2026poster

Neural speech codecs based on Vector-Quantized VAEs (VQ-VAEs) are core audio tokenizers for speech LLMs, yet their reconstruction fidelity is bottlenecked by quantization error. Instead of modifying the quantizer or increasing model capacity—common approaches that complicate downstream language mode…

Cited by 0SourceScholar
2025

AgentStore: Scalable Integration of Heterogeneous Agents As Specialized Generalist Computer Assistant

ACL 2025finding

Digital agents capable of automating complex computer tasks have attracted considerable attention. However, existing agent methods exhibit deficiencies in their generalization and specialization capabilities, especially in handling open-ended computer tasks in real-world environments. Inspired by th…

Cited by 0SourcePDFScholar
2025

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

ICASSP 2025accepted

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control net…

Cited by 0SourceScholar
2025

Binary Representation Learning for Discriminative Acoustic Unit Discovery

ICASSP 2025accepted

Acoustic Unit Discovery (AUD) aims to obtain phoneme-like units that preserve linguistically significant information while removing paralinguistic details. Although Contrastive Predictive Coding (CPC) has emerged as a leading self-supervised representation learning method for this task, CPC-based me…

Cited by 0SourceScholar
2025

Black-Box Adversarial Defense Against Voice Conversion Using Latent Space Perturbation

ICASSP 2025accepted

Voice Conversion (VC) technologies have advanced significantly, enabling voice cloning with just a few seconds of audio, posing serious risks to privacy, property, and reputation. In response to these threats, adversarial defense methods protect users by adding imperceptible perturbations to the aud…

Cited by 0SourceScholar
2025

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

ICASSP 2025accepted

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are limited to deterministic prediction, overlooking the diversi…

Cited by 0SourceScholar
2025

E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech Synthesis

NeurIPS 2025poster

Recent advancements in speech synthesis technology have enriched our daily lives, with high-quality and human-like audio widely adopted across real-world applications. However, malicious exploitation like voice-cloning fraud poses severe security risks. Existing defense techniques struggle to addres…

Cited by 0SourcecodeScholar
2025

Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning

ACL 2025long

Advancing LLM reasoning skills has captivated wide interest. However, current post-training techniques rely heavily on supervisory signals, such as outcome supervision or auxiliary reward models, which face the problem of scalability and high annotation costs. This motivates us to enhance LLM reason…

2025

Identity-Preserving Audio-Driven Holistic Human Motion Video Generation

ICASSP 2025accepted

Generating realistic human motion videos is a pivotal challenge in advancing human-computer interaction. While existing approaches often focus on generating either head or gesture movements from audio, they lack unified control over full-body motion, frequently producing low-resolution and blurred o…

Cited by 0SourceScholar
2025

Implicit Search via Discrete Diffusion: A Study on Chess

ICLR 2025poster

In the post-AlphaGo era, there has been a renewed interest in search techniques such as Monte Carlo Tree Search (MCTS), particularly in their application to Large Language Models (LLMs). This renewed attention is driven by the recognition that current next-token prediction models often lack the abil…

2025

Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language Models

ACL 2025long

One of the primary driving forces contributing to the superior performance of Large Language Models (LLMs) is the extensive availability of human-annotated natural language data, which is used for alignment fine-tuning. This inspired researchers to investigate self-training methods to mitigate the e…

2025

LeVo: High-Quality Song Generation with Multi-Preference Alignment

NeurIPS 2025poster

Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limit…

Cited by 0SourcecodeScholar
2025

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data

ICASSP 2025accepted

Empathetic dialogue is crucial for natural human-computer interaction, allowing the dialogue system to respond in a more personalized and emotionally aware manner, improving user satisfaction and engagement. The emergence of large language models (LLMs) has revolutionized dialogue generation by harn…

Cited by 0SourceScholar
2025

MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement

AAAI 2025technical

Existing works in single-image human reconstruction suffer from weak generalizability due to insufficient training data or 3D inconsistencies for a lack of comprehensive multi-view knowledge. In this paper, we introduce MagicMan, a human-specific multi-view diffusion model to generate high-quality n…

Cited by 8SourcePDFScholar
2025

OS-ATLAS: Foundation Action Model for Generalist GUI Agents

ICLR 2025spotlight

Existing efforts in building GUI agents heavily rely on the availability of robust commercial Vision-Language Models (VLMs) such as GPT-4o and GeminiProVision. Practitioners are often reluctant to use open-source VLMs due to their significant performance lag compared to their closed-source counterpa…

Cited by 29SourcePDFScholar
2025

OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

ACL 2025long

Graphical User Interface (GUI) agents powered by Vision-Language Models (VLMs) have demonstrated human-like computer control capability. Despite their utility in advancing digital automation, the development of such agents faces a critical bottleneck: collecting high-quality trajectory data for trai…

Cited by 0SourcePDFScholar
2025

Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis

ICASSP 2025accepted

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal synchronization adapter, our method significantly improves semantic in…

Cited by 0SourceScholar
2025

Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features

ICASSP 2025accepted

Melody preservation is crucial in singing voice conversion (SVC). However, in many scenarios, audio is often accompanied with background music (BGM), which can cause audio distortion and interfere with the extraction of melody and other key features, significantly degrading SVC performance. Previous…

Cited by 0SourceScholar
2025

VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweening

IJCAI 2025

We propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar

Cited by 0SourcePDFScholar
2025

𝜙-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation

ACL 2025long

Inference-time optimization scales computation to derive deliberate reasoning steps for effective performance. While previous search-based strategies address the short-sightedness of auto-regressive generation, the vast search space leads to excessive exploration and insufficient exploitation. To st…

2024

Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis

EMNLP 2024finding

In recent years, the rapid increase in scientific papers has overwhelmed traditional review mechanisms, resulting in varying quality of publications. Although existing methods have explored the capabilities of Large Language Models (LLMs) for automated scientific reviewing, their generated contents…

2024

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

CVPR 2024poster

Co-speech gestures if presented in the lively form of videos can achieve superior visual effects in human-machine interaction. While previous works mostly generate structural human skeletons resulting in the omission of appearance information we focus on the direct generation of audio-driven co-spee…

2024

Consistent and Relevant: Rethink the Query Embedding in General Sound Separation

ICASSP 2024accepted

The query-based audio separation usually employs specific queries to extract target sources from a mixture of audio signals. Currently, most query-based separation models need additional networks to obtain query embedding. In this way, separation model is optimized to be adapted to the distribution…

Cited by 0SourceScholar
2024

Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models

ICASSP 2024accepted

Audio-driven co-speech human gesture generation has made remarkable advancements recently. However, most previous works only focus on single person audio-driven gesture generation. We aim at solving the problem of conversational co-speech gesture generation that considers multiple participants in a…

Cited by 0SourceScholar
2024

EMO: EARTH MOVER DISTANCE OPTIMIZATION FOR AUTO-REGRESSIVE LANGUAGE MODELING

ICLR 2024poster

Neural language models are probabilistic models of human text. They are predominantly trained using maximum likelihood estimation (MLE), which is equivalent to minimizing the forward cross-entropy between the empirical data distribution and the model distribution. However, various degeneration pheno…

2024

Enhancing Expressiveness in Dance Generation Via Integrating Frequency and Music Style Information

ICASSP 2024accepted

Dance generation, as a branch of human motion generation, has attracted increasing attention. Recently, a few works attempt to enhance dance expressiveness, which includes genre matching, beat alignment, and dance dynamics, from certain aspects. However, the enhancement is quite limited as they lack…

Cited by 0SourceScholar
2024

Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction

ICASSP 2024accepted

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in complex, noisy acoustic environments. To address these challenges,…

Cited by 0SourceScholar
2024

Explore 3D Dance Generation via Reward Model from Automatically-Ranked Demonstrations

AAAI 2024technical

This paper presents an Exploratory 3D Dance generation framework, E3D2, designed to address the exploration capability deficiency in existing music-conditioned 3D dance generation models. Current models often generate monotonous and simplistic dance sequences that misalign with human preferences bec…

Cited by 4SourcePDFScholar
2024

FreeTalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness

ICASSP 2024accepted

Current talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation have designed network structures based on individual gesture datasets, which re…

Cited by 0SourceScholar
2024

Generating Stereophonic Music with Single-Stage Language Models

ICASSP 2024accepted

The recent success of audio language models (LMs) has revolutionized the field of neural music generation. Among all audio LM approaches, MusicGen has demonstrated the success of a single-stage LMs based music generation framework, without needing to train multiple LMs. Despite its promising perform…

Cited by 0SourceScholar
2024

How Vocabulary Sharing Facilitates Multilingualism in LLaMA?

ACL 2024findings

Large Language Models (LLMs), often show strong performance on English tasks, while exhibiting limitations on other languages. What is an LLM’s multilingual capability when it is trained only on certain languages? The underlying mechanism remains unclear. This study endeavors to examine the multilin…

2024

Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts

ICASSP 2024accepted

Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker’s voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language model-based TTS models show zero-shot speaker adaptation cap…

Cited by 0SourceScholar
2024

LLM as Prompter: Low-resource Inductive Reasoning on Arbitrary Knowledge Graphs

ACL 2024findings

Knowledge Graph (KG) inductive reasoning, which aims to infer missing facts from new KGs that are not seen during training, has been widely adopted in various applications. One critical challenge of KG inductive reasoning is handling low-resource scenarios with scarcity in both textual and structura…

2024

Multi-View Midivae: Fusing Track- and Bar-View Representations for Long Multi-Track Symbolic Music Generation

ICASSP 2024accepted

Variational Autoencoders (VAEs) constitute a crucial component of neural symbolic music generation, among which some works have yielded outstanding results and attracted considerable attention. Nevertheless, previous VAEs still encounter issues with overly long feature sequences and generated result…

Cited by 0SourceScholar
2024

Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion

ICASSP 2024accepted

Any-to-any singing voice conversion (SVC) is confronted with the challenge of "timbre leakage" issue caused by inadequate disentanglement between the content and the speaker timbre. To address this issue, this study introduces NeuCoSVC, a novel neural concatenative SVC framework. It consists of a se…

Cited by 0SourceScholar
2024

SCNet: Sparse Compression Network for Music Source Separation

ICASSP 2024accepted

Deep learning-based methods have made significant achievements in music source separation. However, obtaining good results while maintaining a low model complexity remains challenging in super wide-band music source separation. Previous works either overlook the differences in subbands or inadequate…

Cited by 0SourceScholar
2024

SECap: Speech Emotion Captioning with Large Language Model

AAAI 2024technical

Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human spee…

2024

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

ACL 2024long

Graphical User Interface (GUI) agents are designed to automate complex tasks on digital devices, such as smartphones and desktops. Most existing GUI agents interact with the environment through extracted structured data, which can be notably lengthy (e.g., HTML) and occasionally inaccessible (e.g.,…

2024

SimCalib: Graph Neural Network Calibration Based on Similarity between Nodes

AAAI 2024technical

Graph neural networks (GNNs) have exhibited impressive performance in modeling graph data as exemplified in various applications. Recently, the GNN calibration problem has attracted increasing attention, especially in cost-sensitive scenarios. Previous work has gained empirical insights on the issue…

Cited by 6SourcePDFScholar
2024

SongCreator: Lyrics-based Universal Song Generation

NeurIPS 2024poster

Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as singing voice, vocal composition and instrumental arrangement, etc., generating so…

2024

Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech Synthesis

ICASSP 2024accepted

The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-supervised style enhancing method with VQ-VAE-based pre-training for expressive a…

Cited by 0SourceScholar
2024

Symbol-LLM: Towards Foundational Symbol-centric Interface For Large Language Models

ACL 2024long

Although Large Language Models (LLMs) demonstrate remarkable ability in processing and generating human-like text, they do have limitations when it comes to comprehending and expressing world knowledge that extends beyond the boundaries of natural language(e.g., chemical molecular formula). Injectin…

2024

Unifying One-Shot Voice Conversion and Cloning with Disentangled Speech Representations

ICASSP 2024accepted

We propose unifying one-shot voice conversion and cloning into a single model that can be end-to-end optimized. To achieve this, we introduce a novel extension to a speech variational auto-encoder (VAE) that disentangles speech into content and speaker representations. Instead of using a fixed Gauss…

Cited by 0SourceScholar
2023

A Synthetic Corpus Generation Method for Neural Vocoder Training

ICASSP 2023accepted

Nowadays, neural vocoders are preferred for their ability to synthesize high-fidelity audio. However, training a neural vocoder requires a massive corpus of high-quality real audio, and the audio recording process is often labor-intensive. In this work, we propose a synthetic corpus generation metho…

Cited by 0SourceScholar
2023

Av-Sepformer: Cross-Attention Sepformer for Audio-Visual Target Speaker Extraction

ICASSP 2023accepted

Visual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross- and self-attention to fuse and model features from audio an…

Cited by 0SourceScholar
2023

CB-Conformer: Contextual Biasing Conformer for Biased Word Recognition

ICASSP 2023accepted

Due to the mismatch between the source and target domains, how to better utilize the biased word information to improve the performance of the automatic speech recognition model in the target domain becomes a hot research topic. Previous approaches either decode with a fixed external language model…

Cited by 0SourceScholar
2023

Can We Edit Factual Knowledge by In-Context Learning?

EMNLP 2023long main

Previous studies have shown that large language models (LLMs) like GPTs store massive factual knowledge in their parameters. However, the stored knowledge could be false or outdated. Traditional knowledge editing methods refine LLMs via fine-tuning on texts containing specific knowledge. However, wi…

Cited by 0SourcecodeScholar
2023

Compositional Exemplars for In-context Learning

ICML 2023poster

Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task simply by conditioning on a prompt consisting of input-output examples as demonstration, without any parameter updates. The performance of ICL is highly dominat…

2023

Context-Aware Coherent Speaking Style Prediction with Hierarchical Transformers for Audiobook Speech Synthesis

ICASSP 2023accepted

Recent advances in text-to-speech have significantly improved the expressiveness of synthesized speech. However, it is still challenging to generate speech with contextually appropriate and coherent speaking style for multi-sentence text in audiobooks. In this paper, we propose a context-aware coher…

Cited by 0SourceScholar
2023

DASA: Difficulty-Aware Semantic Augmentation for Speaker Verification

ICASSP 2023accepted

Data augmentation is vital to the generalization ability and robustness of deep neural networks (DNNs) models. Existing augmentation methods for speaker verification manipulate the raw signal, which are time-consuming and the augmented samples lack diversity. In this paper, we present a novel diffic…

Cited by 0SourceScholar
2023

DiffuSeq-v2: Bridging Discrete and Continuous Text Spaces for Accelerated Seq2Seq Diffusion Models

EMNLP 2023short findings

Diffusion models have gained prominence in generating high-quality sequences of text. Nevertheless, current approaches predominantly represent discrete text within a continuous diffusion space, which incurs substantial computational overhead during training and results in slower sampling speeds. In…

Cited by 0SourcecodeScholar
2023

DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

ICLR 2023poster

Recently, diffusion models have emerged as a new paradigm for generative models. Despite the success in domains using continuous signals such as vision and audio, adapting diffusion models to natural language is under-explored due to the discrete nature of texts, especially for conditional generatio…

2023

DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models

IJCAI 2023poster

The art of communication beyond speech there are gestures. The automatic co-speech gesture generation draws much attention in computer animation. It is a challenging task due to the diversity of gestures and the difficulty of matching the rhythm and semantics of the gesture to the corresponding spee…

2023

Enhancing the Vocal Range of Single-Speaker Singing Voice Synthesis with Melody-Unsupervised Pre-Training

ICASSP 2023accepted

The single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based on our previous work, this work proposes a melody-unsupervised multi-speaker pretraining method conducted on a multi-sing…

Cited by 0SourceScholar
2023

GTN-Bailando: Genre Consistent long-Term 3D Dance Generation Based on Pre-Trained Genre Token Network

ICASSP 2023accepted

Music-driven 3D dance generation has become an intensive research topic in recent years with great potential for real-world applications. Most existing methods lack the consideration of genre, which results in genre inconsistency in the generated dance movements. In addition, the correlation between…

Cited by 0SourceScholar
2023

Gesper: A Unified Framework for General Speech Restoration

ICASSP 2023accepted

This paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by spee…

Cited by 0SourceScholar
2023

Inter-Subnet: Speech Enhancement with Subband Interaction

ICASSP 2023accepted

Subband-based approaches process subbands in parallel through the model with shared parameters to learn the commonality of local spectrums for noise reduction. In this way, they have achieved remarkable results with fewer parameters. However, in some complex environments, the lack of global spectral…

Cited by 0SourceScholar
2023

Keyword-Specific Acoustic Model Pruning for Open-Vocabulary Keyword Spotting

ICASSP 2023accepted

The open-vocabulary KWS system allows users to customize wake words, but its application is limited by the model size. In this paper, we design a dynamic acoustic model with input-dependent parameters. We find that acoustic frames with similar pronunciation generate similar subnetworks, and differen…

Cited by 0SourceScholar
2023

Lexicon-injected Semantic Parsing for Task-Oriented Dialog

ICASSP 2023accepted

Recently, semantic parsing using hierarchical representations for dialog systems has captured substantial attention. Task-Oriented Parse (TOP), a tree representation with intents and slots as labels of nested tree nodes, has been proposed for parsing user utterances. Previous TOP parsing methods are…

Cited by 0SourceScholar
2023

LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech

ICASSP 2023accepted

Recent advances in neural text-to-speech (TTS) models bring thousands of TTS applications into daily life, where models are deployed in cloud to provide services for customs. Among these models are diffusion probabilistic models (DPMs), which can be stably trained and are more parameter-efficient co…

Cited by 0SourceScholar
2023

QPGesture: Quantization-Based and Phase-Guided Motion Matching for Natural Speech-Driven Gesture Generation

CVPR 2023highlight

Speech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a novel quantization-based and phase-guided motion matching framew…

2023

Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering

ACL 2023long

Despite the surprising few-shot performance of in-context learning (ICL), it is still a common practice to randomly sample examples to serve as context. This paper advocates a new principle for ICL: self-adaptive in-context learning. The self-adaption mechanism is introduced to help each sample find…

2023

Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning

ICLR 2023top-25%

There is a rising interest in further exploring the zero-shot learning potential of large pre-trained language models (PLMs). A new paradigm called data-generation-based zero-shot learning has achieved impressive success. In this paradigm, the synthesized data from the PLM acts as the carrier of kno…

2023

TFCnet: Time-Frequency Domain Corrector for Speech Separation

ICASSP 2023accepted

Deep learning-based methods have made significant achievements in speech separation. Especially the time-domain separation methods have achieved the best performance in recent years. However, time-domain methods are unstable for waveform transformation, which is prone to amplitude and phase errors.…

Cited by 0SourceScholar
2023

TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty

ICASSP 2023accepted

In this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply length penalty (i.e., by trimming trailing frames, see Fig. 1-(b)) directly on the spectrogram of input utterances, which do…

Cited by 0SourceScholar
2023

Unsupervised Explanation Generation via Correct Instantiations

AAAI 2023technical

While large pre-trained language models (PLM) have shown their great skills at solving discriminative tasks, a significant gap remains when compared with humans for explanation-related tasks. Among them, explaining the reason why a statement is wrong (e.g., against commonsense) is incredibly challen…

2023

Wavsyncswap: End-To-End Portrait-Customized Audio-Driven Talking Face Generation

ICASSP 2023accepted

Audio-driven talking face with portrait customization enhances the flexibility of avatar applications for different scenarios, such as on-line meetings, mixed reality, and data generation. Among the existing methods, audio-driven talking face and face swapping are typically viewed as separate tasks…

Cited by 0SourceScholar
2023

What Does Your Face Sound Like? 3D Face Shape towards Voice

AAAI 2023technical

Face-based speech synthesis provides a practical solution to generate voices from human faces. However, directly using 2D face images leads to the problems of uninterpretability and entanglement. In this paper, to address the issues, we introduce 3D face shape which (1) has an anatomical relationshi…

2022

A Character-Level Span-Based Model for Mandarin Prosodic Structure Prediction

ICASSP 2022accepted

The accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation from previous word segmentation results. In this paper, we propose a span-based…

Cited by 0SourceScholar
2022

Adversarial Sample Detection for Speaker Verification by Neural Vocoders

ICASSP 2022accepted

Automatic speaker verification (ASV), one of the most important technology for biometric identification, has been widely adopted in security-critical applications. However, ASV is seriously vulnerable to recently emerged adversarial attacks, yet effective counter-measures against them are limited. I…

Cited by 0SourceScholar
2022

An Approach to Mispronunciation Detection and Diagnosis with Acoustic, Phonetic and Linguistic (APL) Embeddings

ICASSP 2022accepted

Many mispronunciation detection and diagnosis (MD&D) research approaches try to exploit both the acoustic and linguistic features as input. Yet the improvement of the performance is limited, partially due to the shortage of large amount annotated training data at the phoneme level. Phonetic embeddin…

Cited by 0SourceScholar
2022

An End-to-End Chinese Text Normalization Model Based on Rule-Guided Flat-Lattice Transformer

ICASSP 2022accepted

Text normalization, defined as a procedure transforming nonstandard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. Rule-based methods without considering context can not eliminate ambiguation, whereas sequence-to-sequence neural network…

Cited by 0SourceScholar
2022

CoLo: A Contrastive Learning Based Re-ranking Framework for One-Stage Summarization

COLING 2022main

Traditional training paradigms for extractive and abstractive summarization systems always only use token-level or sentence-level training objectives. However, the output summary is always evaluated from summary-level which leads to the inconsistency in training and evaluation. In this paper, we pro…

2022

Disentangling Content and Fine-Grained Prosody Information Via Hybrid ASR Bottleneck Features for Voice Conversion

ICASSP 2022accepted

Non-parallel data voice conversion (VC) have achieved considerable breakthroughs recently through introducing bottleneck features (BNFs) extracted by the automatic speech recognition(ASR) model. However, selection of BNFs have a significant impact on VC result. For example, when extracting BNFs from…

Cited by 0SourceScholar
2022

Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-Based Multi-Modal Context Modeling

ICASSP 2022accepted

Comparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art context modeling methods in conversational TTS only model the textual information in…

Cited by 0SourceScholar
2022

FullSubNet+: Channel Attention Fullsubnet with Complex Spectrograms for Speech Enhancement

ICASSP 2022accepted

Previously proposed FullSubNet has achieved outstanding performance in Deep Noise Suppression (DNS) Challenge and attracted much attention. However, it still encounters issues such as input-output mismatch and coarse processing for frequency bands. In this paper, we propose an extended single-channe…

Cited by 0SourceScholar
2022

Lexical Knowledge Internalization for Neural Dialog Generation

ACL 2022long

We propose knowledge internalization (KI), which aims to complement the lexical knowledge into neural dialog models. Instead of further conditioning the knowledge-grounded dialog (KGD) models on externally retrieved knowledge, we seek to integrate knowledge about each input token internally into the…

2022

Neufa: Neural Network Based End-to-End Forced Alignment with Bidirectional Attention Mechanism

ICASSP 2022accepted

Although deep learning and end-to-end models have been widely used and shown their superiority in automatic speech recognition (ASR) and text-to-speech (TTS) synthesis, state-of-the-art forced alignment (FA) models are still based on hidden Markov model (HMM). HMM has limited view of contextual info…

Cited by 0SourceScholar
2022

ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback

EMNLP 2022finding

Recently, dataset-generation-based zero-shot learning has shown promising results by training a task-specific model with a dataset synthesized from large pre-trained language models (PLMs). The final task-specific model often achieves compatible or even better performance than PLMs under the zero-sh…

2022

Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis

ICASSP 2022accepted

Previous works on expressive speech synthesis mainly focus on current sentence. The context in adjacent sentences is neglected, resulting in inflexible speaking style for the same text, which lacks speech variations. In this paper, we propose a hierarchical framework to model speaking style from con…

Cited by 0SourceScholar
2022

Transformer-S2A: Robust and Efficient Speech-to-Animation

ICASSP 2022accepted

We propose a novel robust and efficient Speech-to-Animation (S2A) approach for synchronized facial animation generation in human-computer interaction. Compared with conventional approaches, the proposed approach utilizes phonetic posteriorgrams (PPGs) of spoken phonemes as input to ensure the cross-…

Cited by 0SourceScholar
2022

Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis

COLING 2022main

Naturalness and expressiveness are crucial for audiobook speech synthesis, but now are limited by the averaged global-scale speaking style representation. In this paper, we propose an unsupervised multi-scale context-sensitive text-to-speech model for audiobooks. A multi-scale hierarchical context e…

Cited by 9SourcePDFScholar
2022

ZeroGen: Efficient Zero-shot Learning via Dataset Generation

EMNLP 2022main

There is a growing interest in dataset generation recently due to the superior generative capacity of large pre-trained language models (PLMs). In this paper, we study a flexible and efficient zero-short learning method, ZeroGen.Given a zero-shot task, we first generate a dataset from scratch using…

2021

Adversarial Defense for Automatic Speaker Verification by Cascaded Self-Supervised Learning Models

ICASSP 2021accepted

Automatic speaker verification (ASV) is one of the core technologies in biometric identification. With the ubiquitous usage of ASV systems in safety-critical applications, more and more malicious attackers attempt to launch adversarial attacks at ASV systems. In the midst of the arms race between at…

Cited by 0SourceScholar
2021

Emotion Controllable Speech Synthesis Using Emotion-Unlabeled Dataset with the Assistance of Cross-Domain Speech Emotion Recognition

ICASSP 2021accepted

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS synthesis on a TTS dataset without emotion labels. Specificall…

Cited by 0SourceScholar
2021

Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation

ACL 2021long

A neural multimodal machine translation (MMT) system is one that aims to perform better translation by extending conventional text-only translation models with multimodal information. Many recent studies report improvements when equipping their models with the multimodal module, despite the controve…

2021

Improving Pronunciation Assessment Via Ordinal Regression with Anchored Reference Samples

ICASSP 2021accepted

Sentence level pronunciation assessment is important for Computer Assisted Language Learning (CALL). Traditional speech pronunciation assessment, based on the Goodness of Pronunciation (GOP) algorithm, has some weakness in assessing a speech utterance: 1) Phoneme GOP scores cannot be easily translat…

Cited by 0SourceScholar
2021

Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach

AAAI 2021technical

Effective emotion inference from user queries helps to give a more personified response for Voice Dialogue Applications(VDAs). The tremendous amounts of VDA users bring in diverse emotion expressions. How to achieve a high emotion inferring performance from large-scale Internet Voice Data in VDAs? T…

Cited by 16SourcePDFScholar
2021

Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language Understanding

EMNLP 2021main

Lack of training data presents a grand challenge to scaling out spoken language understanding (SLU) to low-resource languages. Although various data augmentation approaches have been proposed to synthesize training data in low-resource target languages, the augmented data sets are often noisy, and t…

2021

Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input

ICASSP 2021accepted

Non-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic speech recognition (ASR). Most of the NAR transformers take a fixed-length sequence filled with MASK tokens or a redundan…

Cited by 0SourceScholar
2021

Syntactic Representation Learning For Neural Network Based TTS with Syntactic Parse Tree Traversal

ICASSP 2021accepted

Syntactic structure of a sentence text is correlated with the prosodic structure of the speech that is crucial for improving the prosody and naturalness of a text-to-speech (TTS) system. Nowadays TTS systems usually try to incorporate syntactic structure information with manually designed features b…

Cited by 0SourceScholar
2021

The Huya Multi-Speaker and Multi-Style Speech Synthesis System for M2voc Challenge 2020

ICASSP 2021accepted

Text-to-speech systems now can generate speech that is hard to distinguish from human speech. In this paper, we propose the Huya multi-speaker and multi-style speech synthesis system which is based on DurIAN and HiFi-GAN to generate high-fidelity speech even under low-resource condition. We use the…

Cited by 0SourceScholar
2021

The Multi-Speaker Multi-Style Voice Cloning Challenge 2021

ICASSP 2021accepted

The Multi-speaker Multi-style Voice Cloning Challenge (M2VoC) aims to provide a common sizable dataset as well as a fair testbed for the benchmarking of the popular voice cloning task. Specifically, we formulate the challenge to adapt an average TTS model to the stylistic target voice with limited d…

Cited by 0SourceScholar
2020

Code-Switched Speech Synthesis Using Bilingual Phonetic Posteriorgram with Only Monolingual Corpora

ICASSP 2020accepted

Synthesizing fluent code-switched (CS) speech with consistent voice using only monolingual corpora is still a challenging task, since language alternation seldom occurs during training and the speaker identity is directly correlated with language. In this paper, we present a bilingual phonetic poste…

Cited by 0SourceScholar
2020

End-To-End Accent Conversion Without Using Native Utterances

ICASSP 2020accepted

Techniques for accent conversion (AC) aim to convert non-native to native accented speech. Conventional AC methods try to convert only the speaker identity of a native speaker's voice to that of the non-native accented target speaker, leaving the underlying content and pronunciations unchanged. This…

Cited by 0SourceScholar
2019

A Compact Framework for Voice Conversion Using Wavenet Conditioned on Phonetic Posteriorgrams

ICASSP 2019accepted

Voice conversion can benefit from WaveNet vocoder with improvement in converted speech's naturalness and quality. However, nowadays approaches segregate the training of conversion module and WaveNet vocoder towards different optimization objectives, which might lead to the difficulty in model tuning…

Cited by 0SourceScholar
2019

Dilated Residual Network with Multi-head Self-attention for Speech Emotion Recognition

ICASSP 2019accepted

Speech emotion recognition (SER) plays an important role in intelligent speech interaction. One vital challenge in SER is to extract emotion-relevant features from speech signals. In state-of-the-art SER techniques, deep learning methods, e.g, Convolutional Neural Networks (CNNs), are widely employe…

Cited by 0SourceScholar
2019

End-to-end Code-switched TTS with Mix of Monolingual Recordings

ICASSP 2019accepted

State-of-the-art text-to-speech (TTS) synthesis models can produce monolingual speech with high intelligibility and naturalness. However, when the models are applied to synthesize code-switched (CS) speech, the performance declines seriously. Conventionally, developing a CS TTS system requires multi…

Cited by 0SourceScholar
2019

Learning Discriminative Features from Spectrograms Using Center Loss for Speech Emotion Recognition

ICASSP 2019accepted

Identifying the emotional state from speech is essential for the natural interaction of the machine with the speaker. However, extracting effective features for emotion recognition is difficult, as emotions are ambiguous. We propose a novel approach to learn discriminative features from variable len…

Cited by 0SourceScholar
2019

NN-based Ordinal Regression for Assessing Fluency of ESL Speech

ICASSP 2019accepted

Automatic assessment of a language learner's speech fluency is highly desirable for language education, e.g. for English as a Second Language (ESL) learning. In this paper, we formulate the fluency assessment as a problem of Ordinal Regression with Anchored Reference Samples (ORARS), where the fluen…

Cited by 0SourceScholar
2019

Quasi-fully Convolutional Neural Network with Variational Inference for Speech Synthesis

ICASSP 2019accepted

Recurrent neural networks, such as gated recurrent units (GRUs) and long short-term memory (LSTM), are widely used on acoustic modeling for speech synthesis. However, such sequential generating processes are not friendly to today’s massively parallel computing devices. We introduce a fully convoluti…

Cited by 0SourceScholar
2019

Speech Emotion Recognition Using Capsule Networks

ICASSP 2019accepted

Speech emotion recognition (SER) is a fundamental step towards fluent human-machine interaction. One challenging problem in SER is obtaining utterance-level feature representation for classification. Recent works on SER have made significant progress by using spectrogram features and introducing neu…

Cited by 0SourceScholar
2018

Applying Multitask Learning to Acoustic-Phonemic Model for Mispronunciation Detection and Diagnosis in L2 English Speech

ICASSP 2018accepted

For mispronunciation detection and diagnosis (MDD), nowadays approaches generally treat the phonemes in correct and mispronunciations as the same despite the fact they may actually carry different characteristics. Furthermore, serious data imbalance issue between correct and mispronunciation in data…

Cited by 0SourceScholar
2018

Emphatic Speech Generation with Conditioned Input Layer and Bidirectional LSTMS for Expressive Speech Synthesis

ICASSP 2018accepted

By highlighting the focus of an utterance to draw attention, emphasis in speech interaction plays an important role for speaker intention expressing and understanding. Therefore, emphatic speech synthesis draws increasing interest in the text-to-speech (TTS) area. For emphatic speech synthesis, thre…

Cited by 0SourceScholar
2018

Feature Based Adaptation for Speaking Style Synthesis

ICASSP 2018accepted

Speaking style plays an important role in the expressivity of speech for communication. Hence speaking style is very important for synthetic speech as well. Speaking style adaptation faces the difficulty that the data of specific styles may be limited and difficult to obtain in large amounts. A poss…

Cited by 0SourceScholar
2018

Unsupervised Discovery of an Extended Phoneme Set in L2 English Speech for Mispronunciation Detection and Diagnosis

ICASSP 2018accepted

Second language (L2) speech is often labelled with the native, phoneme categories. Hence, we often observe segments for which it is difficult, if not impossible, to decide on a categorical phoneme label. We refer to these segments as “non-categorical” phoneme units. Existing approaches to mispronunc…

Cited by 0SourceScholar
2017

Learning cross-lingual knowledge with multilingual BLSTM for emphasis detection with limited training data

ICASSP 2017accepted

Bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) has achieved state-of-the-art performance in many sequence processing problems given its capability in capturing contextual information. However, for languages with limited amount of training data, it is still difficult to o…

Cited by 0SourceScholar
2017

Multi-task learning of structured output layer bidirectional LSTMS for speech synthesis

ICASSP 2017accepted

Recurrent neural networks (RNNs) and their bidirectional long short term memory (BLSTM) variants are powerful sequence modelling approaches. Their inherently strong ability in capturing long range temporal dependencies allow BLSTM-RNN speech synthesis systems to produce higher quality and smoother s…

Cited by 0SourceScholar
2016

Learning cross-lingual information with multilingual BLSTM for speech synthesis of low-resource languages

ICASSP 2016accepted

Bidirectional long short-term memory (BLSTM) based speech synthesis has shown great potential in improving the quality of the synthetic speech. However, for low-resource languages, it is difficult to obtain a high quality BLSTM model. BLSTM based speech synthesis can be viewed as a transformation be…

Cited by 0SourceScholar
2016

Low level descriptors based DBLSTM bottleneck feature for speech driven talking avatar

ICASSP 2016accepted

Speech is bimodal in nature. There are close correlations between the acoustic speech signals and the visual gestures such as lip movements, facial expressions and head motions. For speech driven talking avatar, how to derive more representative acoustic features from which to predict more accurate…

Cited by 0SourceScholar
2016

Question detection from acoustic features using recurrent neural network with gated recurrent unit

ICASSP 2016accepted

Question detection is of importance for many speech applications. Only parts of the speech utterances can provide useful clues for question detection. Previous work of question detection using acoustic features in Mandarin conversation is weak in capturing such proper time context information, which…

Cited by 0SourceScholar
2015

A deep recurrent approach for acoustic-to-articulatory inversion

ICASSP 2015accepted

To solve the acoustic-to-articulatory inversion problem, this paper proposes a deep bidirectional long short term memory recurrent neural network and a deep recurrent mixture density network. The articulatory parameters of the current frame may have correlations with the acoustic features many frame…

Cited by 0SourceScholar
2015

HMM-based emphatic speech synthesis for corrective feedback in computer-aided pronunciation training

ICASSP 2015accepted

This paper investigates the incorporation of hidden Markov model (HMM) based emphatic speech synthesis for audio exaggeration into an audio-visual speech synthesis framework for the corrective feedback in computer-aided pronunciation training (CAPT). To improve the voice quality of the synthetic emp…

Cited by 0SourceScholar