← Search

Zhengqi Wen

24 accepted papers

2026

AStar: Boosting Multimodal Reasoning with Automated Structured Thinking

AAAI 2026technical

Multimodal large language models excel across diverse domains but struggle with complex visual reasoning tasks. To enhance their reasoning capabilities, current approaches typically rely on explicit search or post-training techniques. However, search-based methods suffer from computational inefficie

Cited by 0SourcePDFScholar
2026

Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs

ICLR 2026poster

Knowledge distillation has emerged as a pivotal technique for transferring knowledge from stronger large language models (LLMs) to smaller, more efficient models. However, traditional distillation approaches face challenges related to knowledge conflicts and high resource demands, particularly when…

Cited by 0SourceScholar
2026

FAKE SPEECH WILD: DETECTING DEEPFAKE SPEECH ON SOCIAL MEDIA PLATFORM

ICASSP 2026poster

The rapid advancement of speech generation technology has led to the widespread proliferation of deepfake speech across social media platforms. While deepfake audio countermeasures (CMs) achieve promising results on public datasets, their performance degrades significantly in cross-domain scenarios.…

Cited by 0SourcePDFScholar
2026

OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech

ICASSP 2026oral

Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a direct combination of audio-related labels or their diverse rephrasings, making it difficult to handle flexible, high-level…

Cited by 0SourcePDFScholar
2026

PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis

AAAI 2026technical

Multimodal sentiment analysis (MSA) is a research field that recognizes human sentiments by combining textual, visual, and audio modalities. The main challenge lies in integrating sentiment-related information from different modalities, which typically arises during the unimodal feature extraction p

Cited by 0SourcePDFScholar
2025

Code-switching Mediated Sentence-level Semantic Learning

AAAI 2025technical

Code-switching is a linguistic phenomenon in which different languages are used interactively during conversation. It poses significant performance challenges to natural language processing (NLP) tasks due to the often monolingual nature of the underlying system. We focus on sentence-level semantic…

Cited by 0SourcePDFScholar
2025

DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech

ICASSP 2025accepted

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the…

Cited by 0SourceScholar
2025

DetailTTS: Learning Residual Detail Information for Zero-shot Text-to-speech

ICASSP 2025accepted

Traditional text-to-speech (TTS) systems often face challenges in aligning text and speech, leading to the omission of critical linguistic and acoustic details. This misalignment creates an information gap, which existing methods attempt to address by incorporating additional inputs, but these often…

Cited by 0SourceScholar
2025

ImViD: Immersive Volumetric Videos for Enhanced VR Engagement

CVPR 2025highlight

User engagement is greatly enhanced by fully immersive multimodal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture, large 6-DoF interactive space, Multi-modal feedback, and high…

Cited by 0SourcePDFScholar
2025

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

IJCAI 2025

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which h

2025

MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection

ICASSP 2025accepted

Multimodal fake news detection is essential for maintaining the authenticity of Internet multimedia information. Significant differences in form and content of multimodal information lead to intensified optimization conflicts, hindering effective model training as well as reducing the effectiveness…

Cited by 0SourceScholar
2025

Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0

ICASSP 2025accepted

Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the…

Cited by 0SourceScholar
2025

RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing

EMNLP 2025

The rapid advancements in large language models (LLMs) have led to the emergence of routing techniques, which aim to efficiently select the optimal LLM from diverse candidates to tackle specific tasks, optimizing performance while reducing costs. Current LLM routing methods are limited in effectiven

Cited by 0SourcePDFScholar
2023

Learning From Yourself: A Self-Distillation Method For Fake Speech Detection

ICASSP 2023accepted

In this paper, we propose a novel self-distillation method for fake speech detection (FSD), which can significantly improve the performance of FSD without increasing the model complexity. For FSD, some fine-grained information is very important, such as spectrogram defects, mute segments, and so on,…

Cited by 0SourceScholar
2022

ADD 2022: the first Audio Deep Synthesis Detection Challenge

ICASSP 2022accepted

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three t…

Cited by 0SourceScholar
2022

Context-Aware Mask Prediction Network for End-to-End Text-Based Speech Editing

ICASSP 2022accepted

The text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural and it is not obvious how to synthesize records…

Cited by 0SourceScholar
2021

Bi-Level Style and Prosody Decoupling Modeling for Personalized End-to-End Speech Synthesis

ICASSP 2021accepted

End-to-end framework can generate high-quality and high-similarity speech in the personalized speech synthesis task. However, the generalization of out-of-domain texts is still a challenging task. Limited target data leads to unacceptable errors and poor prosody and similarity performance of the syn…

Cited by 0SourceScholar
2021

Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech Recognition

ICASSP 2021accepted

Despite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models’ performance. In this paper, we propose a decoupled transformer model to use mono-lingual paired data and unpaired text…

Cited by 0SourceScholar
2021

Prosody and Voice Factorization for Few-Shot Speaker Adaptation in the Challenge M2voc 2021

ICASSP 2021accepted

The paper describes the CASIA speech synthesis system entry for challenge M2VoC 2021. The low similarity and naturalness of synthesized speech remains a challenging problem for speaker adaptation with few resources. Since the end-to-end acoustic model is too complex to interpret, overfitting will oc…

Cited by 0SourceScholar
2020

Focusing on Attention: Prosody Transfer and Adaptative Optimization Strategy for Multi-Speaker End-to-End Speech Synthesis

ICASSP 2020accepted

End-to-end speech synthesis can generate high-quality synthetic speech and achieve high similarity scores with low-resource adaptation data. However, the generalization of out-domain texts is still a challenging task. The limited adaptation data leads to unacceptable errors and the poor prosody perf…

Cited by 0SourceScholar
2020

Synchronous Transformers for end-to-end Speech Recognition

ICASSP 2020accepted

For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition…

Cited by 0SourceScholar
2019

Phoneme Dependent Speaker Embedding and Model Factorization for Multi-speaker Speech Synthesis and Adaptation

ICASSP 2019accepted

This paper presents an architecture to perform speaker adaption in long short-term memory (LSTM) based Mandarin statistical parametric speech synthesis system. Compared with the conventional methods that focused on using fixed global speaker representations in utterance level for speaker recognition…

Cited by 0SourceScholar
2016

Long short term memory recurrent neural network based encoding method for emotion recognition in video

ICASSP 2016accepted

Human emotion is a temporally dynamic event which can be inferred from both audio and video feature sequences. In this paper we investigate the long short term memory recurrent neural network (LSTM-RNN) based encoding method for category emotion recognition in the video. LSTM-RNN is able to incorpor…

Cited by 0SourceScholar