← Search

Jianhua Tao

59 accepted papers

2026

AStar: Boosting Multimodal Reasoning with Automated Structured Thinking

AAAI 2026technical

Multimodal large language models excel across diverse domains but struggle with complex visual reasoning tasks. To enhance their reasoning capabilities, current approaches typically rely on explicit search or post-training techniques. However, search-based methods suffer from computational inefficie

Cited by 0SourcePDFScholar
2026

EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?

ICLR 2026poster

Descriptive Multimodal Emotion Recognition (DMER) has garnered increasing research attention. Unlike traditional discriminative paradigms that rely on predefined emotion taxonomies, DMER aims to describe human emotional state using free-form natural language, enabling finer-grained and more interpre…

Cited by 0SourcecodeScholar
2026

Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs

ICLR 2026poster

Knowledge distillation has emerged as a pivotal technique for transferring knowledge from stronger large language models (LLMs) to smaller, more efficient models. However, traditional distillation approaches face challenges related to knowledge conflicts and high resource demands, particularly when…

Cited by 0SourceScholar
2026

MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular Images

CVPR 2026

We introduce MetricHMSR (Metric Human Mesh and Scene Recovery), a novel approach for metric human mesh and scene recovery from monocular images. Due to unrealistic assumptions in the camera model and inherent challenges in metric perception, existing approaches struggle to achieve human pose and met

Cited by 0SourcecodeScholar
2026

PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis

AAAI 2026technical

Multimodal sentiment analysis (MSA) is a research field that recognizes human sentiments by combining textual, visual, and audio modalities. The main challenge lies in integrating sentiment-related information from different modalities, which typically arises during the unimodal feature extraction p

Cited by 0SourcePDFScholar
2025

Adversarial Training and Gradient Optimization for Partially Deepfake Audio Localization

ICASSP 2025accepted

Partially deepfake audio localization is important in audio forensics. However, existing localization models for partially deepfake audio face two major challenges: distribution shifts between training and testing data as well as insufficient utilization of information from both manipulated regions…

Cited by 0SourceScholar
2025

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

ICML 2025oral

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suff…

2025

BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement

AAAI 2025technical

Although the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the…

Cited by 0SourcePDFScholar
2025

Code-switching Mediated Sentence-level Semantic Learning

AAAI 2025technical

Code-switching is a linguistic phenomenon in which different languages are used interactively during conversation. It poses significant performance challenges to natural language processing (NLP) tasks due to the often monolingual nature of the underlying system. We focus on sentence-level semantic…

Cited by 0SourcePDFScholar
2025

DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech

ICASSP 2025accepted

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the…

Cited by 0SourceScholar
2025

ImViD: Immersive Volumetric Videos for Enhanced VR Engagement

CVPR 2025highlight

User engagement is greatly enhanced by fully immersive multimodal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture, large 6-DoF interactive space, Multi-modal feedback, and high…

Cited by 0SourcePDFScholar
2025

Listen, Watch, and Learn to Feel: Retrieval-Augmented Emotion Reasoning for Compound Emotion Generation

ACL 2025finding

The ability to comprehend human emotion using multimodal large language models (MLLMs) is essential for advancing human-AI interaction and multimodal sentiment analysis. While psychology theory-based human annotations have contributed to multimodal emotion tasks, the subjective nature of emotional p…

Cited by 0SourcePDFScholar
2025

MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection

ICASSP 2025accepted

Multimodal fake news detection is essential for maintaining the authenticity of Internet multimedia information. Significant differences in form and content of multimodal information lead to intensified optimization conflicts, hindering effective model training as well as reducing the effectiveness…

Cited by 0SourceScholar
2025

Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0

ICASSP 2025accepted

Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the…

Cited by 0SourceScholar
2025

OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition

ICML 2025poster

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-apprai…

2025

PET: High-Frequency Temporal Self-Consistency Learning for Partially Deepfake Audio Localization

ICASSP 2025accepted

Partially deepfake audio attacks have attracted the attention recently, and the demand for locating the manipulation regions of partially deepfake audio arises accordingly. However, existing methods are usually proposed based on frame-level authenticity detection or splicing boundaries detection, ne…

Cited by 0SourceScholar
2025

Pandora’s Box or Aladdin’s Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models

ACL 2025long

Retrieval-Augmented Generation (RAG) has emerged as a crucial method for addressing hallucinations in large language models (LLMs). While recent research has extended RAG models to complex noisy scenarios, these explorations often confine themselves to limited noise types and presuppose that noise i…

2025

RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing

EMNLP 2025

The rapid advancements in large language models (LLMs) have led to the emergence of routing techniques, which aim to efficiently select the optimal LLM from diverse candidates to tackle specific tasks, optimizing performance while reducing costs. Current LLM routing methods are limited in effectiven

Cited by 0SourcePDFScholar
2025

Region-Based Optimization in Continual Learning for Audio Deepfake Detection

AAAI 2025technical

Rapid advancements in speech synthesis and voice conversion bring convenience but also new security risks, creating an urgent need for effective audio deepfake detection. Although current models perform well, their effectiveness diminishes when confronted with the diverse and evolving nature of real…

2025

WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification

ICASSP 2025accepted

Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training pr…

Cited by 0SourceScholar
2024

Bilateral Masking with prompt for Knowledge Graph Completion

NAACL 2024findings

The pre-trained language model (PLM) has achieved significant success in the field of knowledge graph completion (KGC) by effectively modeling entity and relation descriptions. In recent studies, the research in this field has been categorized into methods based on word matching and sentence matchin…

Cited by 1SourcePDFScholar
2024

DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection

NeurIPS 2024poster

At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current AAD algorithms overlook the spatial distribution information…

2024

Fewer-Token Neural Speech Codec with Time-Invariant Codes

ICASSP 2024accepted

Language model based text-to-speech (TTS) models, like VALL-E, have gained attention for their outstanding in-context learning capability in zero-shot scenarios. Neural speech codec is a critical component of these models, which can convert speech into discrete token representations. However, excess…

Cited by 0SourceScholar
2024

Multi-Scale Permutation Entropy for Audio Deepfake Detection

ICASSP 2024accepted

With the widespread application of Automatic Speaker Verification (ASV) technology in security authentication, the threat of fake audio attacks looms as a malicious means compromising system security. In this study, we employ the multi-scale permutation entropy (MPE) in audio deepfake detection, whi…

Cited by 0SourceScholar
2024

NLoPT: N-gram Enhanced Low-Rank Task Adaptive Pre-training for Efficient Language Model Adaption

COLING 2024main

Pre-trained Language Models (PLMs) like BERT have achieved superior performance on different downstream tasks, even when such a model is trained on a general domain. Moreover, recent studies have shown that continued pre-training on task-specific data, known as task adaptive pre-training (TAPT), can…

Cited by 1SourcePDFScholar
2024

Progressive Distillation Based on Masked Generation Feature Method for Knowledge Graph Completion

AAAI 2024technical

In recent years, knowledge graph completion (KGC) models based on pre-trained language model (PLM) have shown promising results. However, the large number of parameters and high computational cost of PLM models pose challenges for their application in downstream tasks. This paper proposes a progress…

2024

Pseudo Labels Regularization for Imbalanced Partial-Label Learning

ICASSP 2024accepted

Partial-label learning (PLL) is an important branch of weakly supervised learning where the single ground truth resides in a set of candidate labels, while the research rarely considers the label imbalance. A recent study for imbalanced PLL propose that the combinatorial challenge of partial-label l…

Cited by 0SourceScholar
2024

What to Remember: Self-Adaptive Continual Learning for Audio Deepfake Detection

AAAI 2024technical

The rapid evolution of speech synthesis and voice conversion has raised substantial concerns due to the potential misuse of such technology, prompting a pressing need for effective audio deepfake detection mechanisms. Existing detection models have shown remarkable success in discriminating known de…

Cited by 29SourcePDFScholar
2023

ALIM: Adjusting Label Importance Mechanism for Noisy Partial Label Learning

NeurIPS 2023poster

Noisy partial label learning (noisy PLL) is an important branch of weakly supervised learning. Unlike PLL where the ground-truth label must conceal in the candidate label set, noisy PLL relaxes this constraint and allows the ground-truth label may not be in the candidate label set. To address this c…

2023

Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio Detection

ICML 2023poster

Current fake audio detection algorithms have achieved promising performances on most datasets. However, their performance may be significantly degraded when dealing with audio of a different dataset. The orthogonal weight modification to overcome catastrophic forgetting does not consider the similar…

2023

GCC-Speaker: Target Speaker Localization with Optimal Speaker-Dependent Weighting in Multi-Speaker Scenarios

ICASSP 2023accepted

Existing noise-robust and reverberant-robust localization algorithms fail to localize the target speaker when interfering speakers are present. In this paper, we address the problem of localizing only the target speaker in multi-speaker scenarios and propose a target speaker localization algorithm,…

Cited by 0SourceScholar
2023

M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis

ICASSP 2023accepted

Conversational text-to-speech (TTS) aims to synthesize speech with proper prosody of reply based on the historical conversation. However, it is still a challenge to comprehensively model the conversation, and a majority of conversational TTS systems only focus on extracting global information and om…

Cited by 0SourceScholar
2023

VRA: Variational Rectified Activation for Out-of-distribution Detection

NeurIPS 2023poster

Out-of-distribution (OOD) detection is critical to building reliable machine learning systems in the open world. Researchers have proposed various strategies to reduce model overconfidence on OOD data. Among them, ReAct is a typical and effective technique to deal with model overconfidence, which tr…

2022

ADD 2022: the first Audio Deep Synthesis Detection Challenge

ICASSP 2022accepted

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three t…

Cited by 0SourceScholar
2022

Automatic Depression Level Assessment from Speech By Long-Term Global Information Embedding

ICASSP 2022accepted

Depression is a serious mood disorder which brings negative effects on people's social activities. Therefore, growing attention has been paid to automatic depression assessment, especially from speech. However, most of the previous work uses hand-crafted features or deep neural network-based feature…

Cited by 0SourceScholar
2022

Context-Aware Mask Prediction Network for End-to-End Text-Based Speech Editing

ICASSP 2022accepted

The text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural and it is not obvious how to synthesize records…

Cited by 0SourceScholar
2022

End-to-End Network Based on Transformer for Automatic Detection of Covid-19

ICASSP 2022accepted

The novel coronavirus disease (COVID-19) was declared a pandemic by the World Health Organization. The cumulative number of deaths is more than 4.8 million. Epidemiology experts concur that mass testing is essential for isolating infected individuals, contact tracing, and slowing the progression of…

Cited by 0SourceScholar
2021

Bi-Level Style and Prosody Decoupling Modeling for Personalized End-to-End Speech Synthesis

ICASSP 2021accepted

End-to-end framework can generate high-quality and high-similarity speech in the personalized speech synthesis task. However, the generalization of out-of-domain texts is still a challenging task. Limited target data leads to unacceptable errors and poor prosody and similarity performance of the syn…

Cited by 0SourceScholar
2021

Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech Recognition

ICASSP 2021accepted

Despite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models’ performance. In this paper, we propose a decoupled transformer model to use mono-lingual paired data and unpaired text…

Cited by 0SourceScholar
2021

Multi-Scale and Multi-Region Facial Discriminative Representation for Automatic Depression Level Prediction

ICASSP 2021accepted

Physiological studies have shown that differences in facial activities between depressed patients and normal individuals are manifested in different local facial regions and the durations of these activities are not the same. But most previous works extract features from the entire facial region at…

Cited by 0SourceScholar
2021

Multimodal Cross- and Self-Attention Network for Speech Emotion Recognition

ICASSP 2021accepted

Speech Emotion Recognition (SER) requires a thorough understanding of both the linguistic content of an utterance (i.e., textual information) and how the speaker utters it (i.e., acoustic information). The one vital challenge in SER is how to effectively fuse these two kinds of information. In this…

Cited by 0SourceScholar
2021

Patnet : A Phoneme-Level Autoregressive Transformer Network for Speech Synthesis

ICASSP 2021accepted

Aiming at efficiently predicting acoustic features with high naturalness and robustness, this paper proposes PATNet, a neural acoustic model for speech synthesis using phoneme-level autoregression. PATNet accepts phoneme sequences as input and is built based on Transformer structure. PATNet adopts a…

Cited by 0SourceScholar
2021

Prosody and Voice Factorization for Few-Shot Speaker Adaptation in the Challenge M2voc 2021

ICASSP 2021accepted

The paper describes the CASIA speech synthesis system entry for challenge M2VoC 2021. The low similarity and naturalness of synthesized speech remains a challenging problem for speaker adaptation with few resources. Since the end-to-end acoustic model is too complex to interpret, overfitting will oc…

Cited by 0SourceScholar
2020

Focusing on Attention: Prosody Transfer and Adaptative Optimization Strategy for Multi-Speaker End-to-End Speech Synthesis

ICASSP 2020accepted

End-to-end speech synthesis can generate high-quality synthetic speech and achieve high similarity scores with low-resource adaptation data. However, the generalization of out-domain texts is still a challenging task. The limited adaptation data leads to unacceptable errors and the poor prosody perf…

Cited by 0SourceScholar
2020

Multimodal Transformer Fusion for Continuous Emotion Recognition

ICASSP 2020accepted

Multimodal fusion increases the performance of emotion recognition because of the complementarity of different modalities. Compared with decision level and feature level fusion, model level fusion makes better use of the advantages of deep neural networks. In this work, we utilize the Transformer mo…

Cited by 0SourceScholar
2020

Synchronous Transformers for end-to-end Speech Recognition

ICASSP 2020accepted

For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition…

Cited by 0SourceScholar
2019

Discriminative Video Representation with Temporal Order for Micro-expression Recognition

ICASSP 2019accepted

Micro-expression recognition is a challenging task due to its low intensity and short duration and how to extract the subtle facial changes is a key issue in this field. Although there are many methods attempt to cope with this problem, they are difficult to encode the temporal order of all frames i…

Cited by 0SourceScholar
2019

Language-invariant Bottleneck Features from Adversarial End-to-end Acoustic Models for Low Resource Speech Recognition

ICASSP 2019accepted

This paper proposes to learn language-invariant bottleneck features from an adversarial end-to-end acoustic model for low resource languages. The multilingual end-to-end model is trained with a connectionist temporal classification loss function. The model has shared and private layers. The shared l…

Cited by 0SourceScholar
2019

Phoneme Dependent Speaker Embedding and Model Factorization for Multi-speaker Speech Synthesis and Adaptation

ICASSP 2019accepted

This paper presents an architecture to perform speaker adaption in long short-term memory (LSTM) based Mandarin statistical parametric speech synthesis system. Compared with the conventional methods that focused on using fixed global speaker representations in utterance level for speaker recognition…

Cited by 0SourceScholar
2018

End-to-End Continuous Emotion Recognition from Video Using 3D Convlstm Networks

ICASSP 2018accepted

Conventional continuous emotion recognition consists of feature extraction step followed by regression step. However, the objective of the two steps is not consistent as they are parted. Besides, there is still no consensus about appropriate emotional features. In this study, we propose an end-to-en…

Cited by 0SourceScholar
2017

A novel pitch extraction based on jointly trained deep BLSTM Recurrent Neural Networks with bottleneck features

ICASSP 2017accepted

Pitch is an important characteristic of speech and is useful for many applications. However, it is still challenging to estimate pitch in strong noise. In this paper, we propose a joint training approach to determinate pitch. First, a Bidirectional Long Short-Term Memory Recurrent Neural Networks (B…

Cited by 0SourceScholar
2016

Extraction of tongue contour in real-time magnetic resonance imaging sequences

ICASSP 2016accepted

Real-time magnetic resonance imaging (rtMRI) is becoming a practical tool in speech production research and language pathology observation. It is still a challenge to extract the tongue contour accurately in rtMRI sequences, since tongue is a soft tissue and often touches other organs such as lips a…

Cited by 0SourceScholar
2016

Long short term memory recurrent neural network based encoding method for emotion recognition in video

ICASSP 2016accepted

Human emotion is a temporally dynamic event which can be inferred from both audio and video feature sequences. In this paper we investigate the long short term memory recurrent neural network (LSTM-RNN) based encoding method for category emotion recognition in the video. LSTM-RNN is able to incorpor…

Cited by 0SourceScholar
2015

Estimate articulatory MRI series from acoustic signal using deep architecture

ICASSP 2015accepted

This paper presents our work on acoustic-to-articulatory inversion mapping, in which, the articulatory data is the MRI series for articulators on mid-sagittal plan. Deep architectures based on restricted Boltzmann machine (RBM) and linear regression are employed to construct the audio-visual mapping…

Cited by 0SourceScholar
2015

Evaluation of linear regression for speaker adaptation in HMM-based articulatory movements estimation

ICASSP 2015accepted

Acoustic-to-articulatory inversion problem is usually studied in speaker-specific manner because both articulatory data and acoustic features contain speaker-specific components. This paper presents our work on speaker-adaptation training for this problem. We implement speaker adaptation in HMM-base…

Cited by 0SourceScholar