← Search

Jianwu Dang

51 accepted papers

2026

BREAKING DATA EFFICIENCY DILEMMA: A FEDERATED AND AUGMENTED LEARNING FRAMEWORK FOR ALZHEIMER’S DISEASE DETECTION VIA SPEECH

ICASSP 2026poster

Early diagnosis of Alzheimer's Disease (AD) is crucial for delaying its progression. While AI-based speech detection is non-invasive and cost-effective, it faces a critical data efficiency dilemma due to medical data scarcity and privacy barriers. Therefore, we propose FAL-AD, a novel framework that…

Cited by 0SourcePDFScholar
2025

A Chinese Expressive Long-dialogue Speech Dataset with Scripts

ICASSP 2025accepted

With the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervis…

Cited by 0SourceScholar
2025

A Prompt Learning Framework with Large Language Model Augmentation for Few-shot Multi-label Intent Detection

ICASSP 2025accepted

Intent detection (ID) is essential in spoken language understanding, especially in multi-label settings where intent labels are interdependent and diverse. Existing methods like SE-MLP and QA-FT struggle in few-shot settings, due to limited data availability and efficiency concerns. To address this,…

Cited by 0SourceScholar
2025

Augmenting Short Enrollment Speech via Synthesis for Target Speaker Extraction

ICASSP 2025accepted

A high-quality enrollment speech is crucial to target speaker extraction (TSE), since it provides essential cues for identifying the target speaker in the mixture. However, real applications usually only permit a short enrollment speech, e.g. a wakeup word for a mobile device, that provides limited…

Cited by 0SourceScholar
2025

Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 Challenge

ICASSP 2025accepted

In this paper, we present the system developed by our team, CCATTS, for the LIMMITS’25 challenge, focusing on few-shot and zero-shot TTS. We adopt a two-stage TTS strategy. In track 1, we fine-tune the pre-trained ZMM-TTS model and successfully achieve multilingual low-latency TTS. In track 2, we pr…

Cited by 0SourceScholar
2025

Enriching Multimodal Sentiment Analysis Through Textual Emotional Descriptions of Visual-Audio Content

AAAI 2025technical

Multimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within audio and video expressions poses a formidable challenge, particularly when emot…

2025

HeterGP: Bridging Heterogeneity in Graph Neural Networks with Multi-View Prompting

AAAI 2025technical

The challenges tied to unstructured graph data are manifold, primarily falling into node, edge, and graph-level problem categories. Graph Neural Networks (GNNs) serve as effective tools to tackle these issues. However, individual tasks often demand distinct model architectures, and training these mo…

Cited by 0SourcePDFScholar
2025

Integration of Old and New Knowledge for Generalized Intent Discovery: A Consistency-driven Prototype-Prompting Framework

IJCAI 2025

Intent detection aims to identify user intents from natural language inputs, where supervised methods rely heavily on labeled in-domain (IND) data and struggle with out-of-domain (OOD) intents, limiting their practical applicability. Generalized Intent Discovery (GID) addresses this by leveraging un

2025

Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement

ICASSP 2025accepted

In recent speech enhancement (SE) research, transformer and its variants have emerged as the predominant methodologies. However, the quadratic complexity of the self-attention mechanism imposes certain limitations on practical deployment. Mamba, as a novel state-space model (SSM), has gained widespr…

Cited by 0SourceScholar
2025

Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module

ICASSP 2025accepted

The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coef…

Cited by 5SourceScholar
2025

Rethinking Contrastive Learning in Graph Anomaly Detection: A Clean-View Perspective

IJCAI 2025

Graph anomaly detection aims to identify unusual patterns in graph-based data, with wide applications in fields such as web security and financial fraud detection. Existing methods typically rely on contrastive learning, assuming that a lower similarity between a node and its local subgraph indicate

Cited by 0SourcePDFScholar
2025

Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis

NeurIPS 2025spotlight

While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modelin…

Cited by 0SourceScholar
2024

Ahpatron: A New Budgeted Online Kernel Learning Machine with Tighter Mistake Bound

AAAI 2024technical

In this paper, we study the mistake bound of online kernel learning on a budget. We propose a new budgeted online kernel learning model, called Ahpatron, which significantly improves the mistake bound of previous work and resolves an open problem related to upper bounds of hypothesis space constrain…

2024

High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models

ICASSP 2024accepted

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech representations(semantic & acoustic) and using two sequence-to-seque…

Cited by 0SourceScholar
2024

Learning Speech Representation from Contrastive Token-Acoustic Pretraining

ICASSP 2024accepted

For fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech should serve as a "bridge" between text and acoustic information, containing info…

Cited by 0SourceScholar
2024

Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding

ICASSP 2024accepted

Recently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks to decouple TTS. However, existing methods suffer from three problems: the high-…

Cited by 0SourceScholar
2023

Augmenting Affective Dependency Graph via Iterative Incongruity Graph Learning for Sarcasm Detection

AAAI 2023technical

Recently, progress has been made towards improving automatic sarcasm detection in computer science. Among existing models, manually constructing static graphs for texts and then using graph neural networks (GNNs) is one of the most effective approaches for drawing long-range incongruity patterns. Ho…

Cited by 24SourcePDFScholar
2023

Brain Network Features Differentiate Intentions from Different Emotional Expressions of the Same Text

ICASSP 2023accepted

Intent differentiation in speech communication relies not only on linguistic information but also on paralinguistic information. The same textual content, when pronounced with different prosodies and emotions, may express totally different intentions. The true intentions in this condition can be eas…

Cited by 0SourceScholar
2023

Commonsense Knowledge Enhanced Sentiment Dependency Graph for Sarcasm Detection

IJCAI 2023poster

Sarcasm is widely utilized on social media platforms such as Twitter and Reddit. Sarcasm detection is required for analyzing people's true feelings since sarcasm is commonly used to portray a reversed emotion opposing the literal meaning. The syntactic structure is the key to make better use of comm…

Cited by 16SourcePDFScholar
2023

Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker Verification

ICASSP 2023accepted

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The primary motivation of our cross-modal co-learning method is mo…

Cited by 0SourceScholar
2023

Leveraging Positional-Related Local-Global Dependency for Synthetic Speech Detection

ICASSP 2023accepted

Automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. As synthetic speech exhibits local and global artifacts compared to natural speech, incorporating local-global dependency would lead to better anti-spoofing performance. To this end, we propose the Rawformer that levera…

Cited by 0SourceScholar
2023

Noise-Disentanglement Metric Learning for Robust Speaker Verification

ICASSP 2023accepted

Automatic speaker verification (ASV) suffers from performance degradation in noisy environments. To solve this problem, we propose the noise-disentanglement metric learning to reduce the speaker-irrelevant noisy components and build a noise-invariant embedding space. Specifically, the disentanglemen…

Cited by 0SourceScholar
2023

Self-Supervised Audio-Visual Speaker Representation with Co-Meta Learning

ICASSP 2023accepted

In self-supervised speaker verification, the quality of pseudo labels determines the upper bound of its performance and it is not uncommon to end up with massive amount of unreliable pseudo labels. We observe that the complementary information in different modalities ensures a robust supervisory sig…

Cited by 0SourceScholar
2023

Speech and Noise Dual-Stream Spectrogram Refine Network With Speech Distortion Loss For Robust Speech Recognition

ICASSP 2023accepted

In recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. Traditional joint training methods only use enhanced speech as input for the backend. However, it is difficult for speech…

Cited by 0SourceScholar
2023

Time-Domain Speech Enhancement Assisted by Multi-Resolution Frequency Encoder and Decoder

ICASSP 2023accepted

Time-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS [1] introduces multi-resolution STFT loss to enhance performance. However, some resolutions used for STFT contain non-stationary signals, and it is challenging to learn multi-resolution frequen…

Cited by 0SourceScholar
2023

VF-Taco2: Towards Fast and Lightweight Synthesis for Autoregressive Models with Variation Autoencoder and Feature Distillation

ICASSP 2023accepted

With the development of deep learning, end-to-end neural text-to-speech (TTS) systems have achieved significant improvements in high-quality speech synthesis. However, most of these systems are attention-based autoregressive models, resulting in slow synthesis speed and large model parameter sizes.…

Cited by 0SourceScholar
2022

Cache: Modeling Contribution-Aware Context Hierarchically for Long-Range Dialogue State Tracking

ICASSP 2022accepted

Recently, many studies on dialogue state tracking (DST) based on the copy-augmented encoder-decoder framework have been proposed and have achieved encouraging performance. However, these studies commonly lose earlier information during encoding the long dialogues with RNNs, and have difficulty for t…

Cited by 0SourceScholar
2022

Compressing Transformer-Based ASR Model by Task-Driven Loss and Attention-Based Multi-Level Feature Distillation

ICASSP 2022accepted

The current popular knowledge distillation (KD) methods effectively compress the transformer-based end-to-end speech recognition model. However, existing methods fail to utilize complete information of the teacher model, and they distill only a limited number of blocks of the teacher model. In this…

Cited by 0SourceScholar
2022

Domain-Invariant Feature Learning for Cross Corpus Speech Emotion Recognition

ICASSP 2022accepted

To deal with speech emotion recognition (SER) in real-life applications, researchers have to focus on cross corpus SER, where the feature distribution of source and target datasets are different. In this paper, we propose an efficient domain adversarial training method to cope with the non-affective…

Cited by 0SourceScholar
2022

L-SpEx: Localized Target Speaker Extraction

ICASSP 2022accepted

Speaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s…

Cited by 0SourceScholar
2022

Learning Domain-Invariant Transformation for Speaker Verification

ICASSP 2022accepted

Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors such as recording device and speaking style in real-world applications, which leads to unsatisfactory performance. To this end, we propose the meta generalized transformation via meta-le…

Cited by 0SourceScholar
2022

Multi-Stage Graph Representation Learning for Dialogue-Level Speech Emotion Recognition

ICASSP 2022accepted

With the development of speech emotion recognition (SER), most of current research is utterance-level and cannot fit the need of actual scenarios. In this paper, we propose a novel strategy that focuses on capturing dialogue-level contextual information. On the basis of utterance-level representatio…

Cited by 0SourceScholar
2022

Using Multiple Reference Audios and Style Embedding Constraints for Speech Synthesis

ICASSP 2022accepted

The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding must be manually selected during inference. Due to the fact t…

Cited by 0SourceScholar
2021

Domain-Adversarial Autoencoder with Attention Based Feature Level Fusion for Speech Emotion Recognition

ICASSP 2021accepted

Over the past two decades, although speech emotion recognition (SER) has garnered considerable attention, the problem of insufficient training data has been unresolved. A potential solution for this problem is to pre-train a model and transfer knowledge from large amounts of audio data. However, the…

Cited by 0SourceScholar
2021

Improving Naturalness and Controllability of Sequence-to-Sequence Speech Synthesis by Learning Local Prosody Representations

ICASSP 2021accepted

State-of-the-art neural text-to-speech (TTS) networks are trained with a large amount of speech data, which significantly improves the quality of synthetic speech compared with traditional approaches. However, the prosody and controllability of the generated speech is still insufficient, especially…

Cited by 0SourceScholar
2021

Meta-Learning for Cross-Channel Speaker Verification

ICASSP 2021accepted

Automatic speaker verification (ASV) has been successfully deployed for identity recognition. With increasing use of ASV technology in real-world applications, channel mismatch caused by the recording devices and environments severely degrade its performance, especially in the case of unseen channel…

Cited by 0SourceScholar
2021

Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference Signals

ICASSP 2021accepted

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted s…

Cited by 0SourceScholar
2021

Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation Fusion

ICASSP 2021accepted

Due to the more robust characteristics compared to unimodal, audio-video multimodal emotion recognition (MER) has attracted a lot of attention. The efficiency of representation fusion algorithm often determines the performance of MER. Although there are many fusion algorithms, information redundancy…

Cited by 0SourceScholar
2021

Replay-Attack Detection Using Features With Adaptive Spectro-Temporal Resolution

ICASSP 2021accepted

Variable-resolution processing aims to improve the feature representation ability by enlarging the local discriminative details. In previous anti-spoofing studies, different phones and frequency regions were both proven to have various levels of sensitivity to replay distortion. In this paper, an ad…

Cited by 0SourceScholar
2021

Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion Recognition

ICASSP 2021accepted

Convolutional neural network (CNN) is found to be effective in learning representation for speech emotion recognition. CNNs do not explicitly model the associations or relative importance of features in the spectral/temporal/channel-wise axes. In this paper, we propose an attention module, named spe…

Cited by 0SourceScholar
2021

Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural Network

ICASSP 2021accepted

Voice activity detection (VAD) based on deep learning has achieved remarkable success. However, when the traditional features (e.g., raw waveforms and MFCCs) are directly fed to the deep neural network model, the performance decreases because of noise interference. Here, we propose a robust VAD appr…

Cited by 0SourceScholar
2020

A Hierarchical Model for Dialog Act Recognition Considering Acoustic and Lexical Context Information

ICASSP 2020accepted

Dialog act recognition (DAR) is important to capture speakers' intention in a dialog system. Traditional methods commonly use the lexical information from transcripts, acoustic information from speech, and dialog context information to do DAR. However, in these methods, textual context information m…

Cited by 0SourceScholar
2020

End-to-End Articulatory Modeling for Dysarthric Articulatory Attribute Detection

ICASSP 2020accepted

In this study, we focus on detecting articulatory attribute errors for dysarthric patients with cerebral palsy (CP) or amyotrophic lateral sclerosis (ALS). There are two major challenges for this task. The pronunciation of dysarthric patients is unclear and inaccurate, which results in poor performa…

Cited by 0SourceScholar
2020

Spectrograms Fusion with Minimum Difference Masks Estimation for Monaural Speech Dereverberation

ICASSP 2020accepted

Spectrograms fusion is an effective method for incorporating complementary speech dereverberation systems. Previous linear spectrograms fusion by averaging multiple spectrograms shows outstanding performance. However, various systems with different features cannot apply this simple method. In this s…

Cited by 0SourceScholar
2020

Speech Emotion Recognition with Local-Global Aware Deep Representation Learning

ICASSP 2020accepted

Convolutional neural network (CNN) based deep representation learning methods for speech emotion recognition (SER) have demonstrated great success. The basic design of CNN restricts the ability to model only local information well. Capsule network (CapsNet) can overcome the shortages of CNNs to capt…

Cited by 0SourceScholar
2019

Replay Attack Detection Using Magnitude and Phase Information with Attention-based Adaptive Filters

ICASSP 2019accepted

Automatic Speech Verification (ASV) systems are highly vulnerable to spoofing attacks, and replay attack poses the greatest threat among various spoofing attacks. In this paper, we propose a novel multi-channel feature extraction method with attention-based adaptive filters (AAF). Original phase inf…

Cited by 0SourceScholar
2018

A Feature Fusion Method Based on Extreme Learning Machine for Speech Emotion Recognition

ICASSP 2018accepted

Speech emotion recognition is important to understand users' intention in human-computer interaction. However, it is a challenging task partly because we cannot clearly know which feature and model are effective to distinguish emotions. Previous studies utilize convolutional neural network (CNN) dir…

Cited by 0SourceScholar
2016

Investigations into vowel and consonant structures in articulatory and auditory spaces using Laplacian eigenmaps

ICASSP 2016accepted

Many studies have investigated the relationship between the articulatory and auditory features for isolated speech sound and vowels. For fully understanding the mechanisms of speech production and perception, it is necessary to investigate the consonants in the same way. For this reason, in this stu…

Cited by 0SourceScholar
2015

Vocal responses to frequency modulated composite sinewaves via auditory and vibrotactile pathways

ICASSP 2015accepted

Feedback control mechanisms for speaking have been examined using the transformed auditory feedback (TAF) technique. Previous studies have shown that speakers demonstrate fundamental frequency (F0) changes when they monitor their voice with artificial alterations of F0. However, those studies undere…

Cited by 0SourceScholar