← Search

Zixing Zhang

36 accepted papers

2026

AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic

ICML 2026poster

Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. This process requires accurate visual-to-symbolic construction of circuit structure from images and correct multi-step symbolic derivation, both of which…

Cited by 0SourceScholar
2026

MindTracker: Unveiling Implicit Emotions in Long-Horizon Dialogues

IJCAI 2026

Affective computing has achieved notable success in recognizing explicit emotions from short, isolated dialogue segments. However, human emotions are often implicitly expressed, internally regulated, and dynamically evolve over extended interactions. Existing models struggle to disentangle internal

Cited by 0Scholar
2026

PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network

AAAI 2026technical

Existing multimodal representation learning approaches often rely on simple feature concatenation or unified transformations, which fail to effectively disentangle and leverage common and private information across different modalities in a progressive manner. Moreover, they typically lack adaptive

Cited by 0SourcePDFScholar
2026

SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs

AAAI 2026technical

Neural Speech Codecs face a fundamental trade-off at low bitrates: preserving acoustic fidelity often compromises semantic richness. To address this, we introduce SACodec, a novel codec built upon an asymmetric dual-quantizer that employs our proposed Semantic Anchoring mechanism. This design strate

Cited by 0SourcePDFScholar
2026

Weaving Graph over Tokens: Contextualizing Structured Sequences for LLMs

ICML 2026poster

Generative Graph Language Models (GLMs) must reconcile topology with causal language modeling. Linearization obscures multi-hop connectivity, while encoder-based methods bottleneck token-level reasoning during generation. Viewing context modeling as a form of message passing, we introduce **Weaver**…

Cited by 0SourceScholar
2025

DSSM: Dual State Space Model For Human Motions Generation

ICASSP 2025accepted

Text-driven human motion generation has attracted considerable critical attention in recent years. The task requires generating movements that are diverse, natural, and comfortable in accordance with the text description. However, while generating the human motion, there is a significant gap in the…

Cited by 0SourceScholar
2025

Dual-View Learning for Conversational Emotion Recognition Through Context and Emotion-Shift Modeling

AAAI 2025technical

Conversational Emotion Recognition (CER) has recently been explored through conversational context modeling to learn the emotion distribution, i.e., the likelihood over emotion categories associated with each utterance. While these methods have shown promising results in emotion classification, they…

Cited by 0SourcePDFScholar
2025

Enhanced Multimodal Emotion Recognition in Conversations via Contextual Filtering and Multi-Frequency Graph Propagation

ICASSP 2025accepted

Multimodal Emotion Recognition in Conversations (ERC) plays a crucial role in understanding human language and behavior in real-world scenarios. However, existing research tends to simply concatenate multimodal representations, failing to capture the complex relationships between modalities. Recent…

Cited by 0SourceScholar
2025

GateM2Former: Gated Feature Selection and Expert Modeling in Multimodal Emotion Recognition

ICASSP 2025accepted

In recent years, multimodal emotion recognition (MER) has gained significant attention due to its potential to integrate information from diverse signals. However, existing methods often struggle to effectively capture complex interactions and contextual information both inter- and intra-modalities,…

Cited by 0SourceScholar
2025

MHSDB: A Comprehensive Benchmark for Multimodal Humor and Sarcasm Detection Leveraging Foundation Models

ICASSP 2025accepted

Understanding multimodal humor and sarcasm detection remains a key challenge in artificial intelligence. Despite recent advances, inconsistencies in feature extraction, evaluation methods, and experimental setups have hindered fair comparisons across different approaches. To address this issue, we p…

Cited by 0SourceScholar
2025

Parameter-Efficient Federal-Tuning Enhances Privacy Preserving for Speech Emotion Recognition

ICASSP 2025accepted

The Pre-trained Speech Models (PSMs) generate universal speech representations using self-supervised or weakly-supervised learning from large-scale datasets. It achieves promising performance when fine-tuned for specific tasks such as Speech Emotion Recognition (SER). However, fine-tuning on various…

Cited by 0SourceScholar
2025

ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis

AAAI 2025technical

Prosody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks when synthesizing long sentences with complex structures but also produce unnat…

2025

Rethinking Removal Attack and Fingerprinting Defense for Model Intellectual Property Protection: A Frequency Perspective

IJCAI 2025

Training deep neural networks is resource-intensive, making it crucial to protect their intellectual property from infringement. However, current model ownership resolution (MOR) methods predominantly address general removal attacks that involve weight modifications, with limited research considerin

2025

SSE: A Speaking Style Extractor Based on Fine-Grained Contrastive Learning between Speech and Descriptive Text

ICASSP 2025accepted

Effective extraction of paralinguistic features from speech, such as emotion, accent, and age, remains a challenging task in speech processing. Traditional methods typically address each type of paralinguistic information with separate classification or regression tasks—e. g., emotion recognition, a…

Cited by 0SourceScholar
2025

Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling

ICASSP 2025accepted

The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling met…

Cited by 0SourceScholar
2025

XDGesture: An xLSTM-based Diffusion Model for Co-speech Gesture Generation

ICASSP 2025accepted

In multimodal human-computer interaction, generating co-speech gestures is crucial for enhancing interaction naturalness and user experience. However, achieving synchronized and natural gesture sequences remains a significant challenge due to the complexity of modeling temporal dependencies across d…

Cited by 0SourceScholar
2024

Customising General Large Language Models for Specialised Emotion Recognition Tasks

ICASSP 2024accepted

The advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tasks but also in emotion recognition in terms of accuracy, universality, explanation, robustness, few/zero-shot learning,…

Cited by 0SourceScholar
2024

EmoTransKG: An Innovative Emotion Knowledge Graph to Reveal Emotion Transformation

ACL 2024findings

This paper introduces EmoTransKG, an innovative Emotion Knowledge Graph (EKG) that establishes connections and transformations between emotions across diverse open-textual events. Compared to existing EKGs, which primarily focus on linking emotion keywords to related terms or on assigning sentiment…

2024

Esihgnn: Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion Recognition

ICASSP 2024accepted

Conversational Emotion Recognition (CER) aims to predict the emotion expressed by an utterance (referred to as an "event") during a conversation. Existing graph-based methods mainly focus on event interactions to comprehend the conversational context, while overlooking the direct influence of the sp…

Cited by 0SourceScholar
2024

HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech

ICASSP 2024accepted

Automatically detecting Alzheimer’s Disease (AD) from spontaneous speech plays an important role in its early diagnosis. Recent approaches highly rely on the Transformer architectures due to its efficiency in modelling long-range context dependencies. However, the quadratic increase in computational…

Cited by 0SourceScholar
2024

Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty Estimation

ICASSP 2024accepted

Heart murmurs are a common manifestation of cardiovascular diseases and can provide crucial clues to early cardiac abnormalities. While most current research methods primarily focus on the accuracy of models, they often overlook other important aspects such as the interpretability of machine learnin…

Cited by 0SourceScholar
2024

LSTDial: Enhancing Dialogue Generation via Long- and Short-Term Measurement Feedback

NAACL 2024long

Generating high-quality responses is a key challenge for any open domain dialogue systems. However, even though there exist a variety of quality dimensions especially designed for dialogue evaluation (e.g., coherence and diversity scores), current dialogue systems rarely utilize them to guide the re…

2023

Privacy-Enhanced Federated Learning Against Attribute Inference Attack for Speech Emotion Recognition

ICASSP 2023accepted

Federal learning-based (FL) Speech Emotion Recognition (SER) framework aims to protect data privacy when characterizing emotions. However, previous studies have shown that the framework is vulnerable, because curious servers can indirectly infer user private information. To address this challenge, w…

Cited by 0SourceScholar
2023

Zero-Shot Speech Emotion Recognition Using Generative Learning with Reconstructed Prototypes

ICASSP 2023accepted

Zero-shot Speech Emotion Recognition (SER) enables machines to perceive unseen-emotional speech without knowing any samples from these emotional states, which is helpful in audio-based autonomous affective computing. However, existing works on zero-shot SER directly employ original prototypes and on…

Cited by 7SourceScholar
2022

Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional Network

ICASSP 2022accepted

Automated classification of respiratory sounds has become an active research area in recent years. While recent studies have utilised deep learning methods to aid with respiratory sound classification, the performance is heavily influenced by the datasets available for respiratory sound classificati…

Cited by 0SourceScholar
2020

Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition Models

ICASSP 2020accepted

The development of deep learning models for speech emotion recognition has become a popular area of research. Adversarially generated data can cause false predictions, and in an endeavor to ensure model robustness, defense methods against such attacks should be addressed. With this in mind, in this…

Cited by 0SourceScholar
2020

Hierarchical Attention Transfer Networks for Depression Assessment from Speech

ICASSP 2020accepted

A growing area of mental health research is the search for speech-based objective markers for conditions such as depression. However, when combined with machine learning, this search can be challenging due to a limited amount of annotated training data. In this paper, we propose a novel crosstask ap…

Cited by 0SourceScholar
2019

Attention-augmented End-to-end Multi-task Learning for Emotion Prediction from Speech

ICASSP 2019accepted

Despite the increasing research interest in end-to-end learning systems for speech emotion recognition, conventional systems either suffer from the overfitting due in part to the limited training data, or do not explicitly consider the different contributions of automatically learnt representations…

Cited by 0SourceScholar
2019

Compact Convolutional Recurrent Neural Networks via Binarization for Speech Emotion Recognition

ICASSP 2019accepted

Despite the great advances, most of the recently developed automatic speech recognition systems focus on working in a server-client manner, and thus often require a high computational cost, such as the storage size and memory accesses. This, however, does not satisfy the increasing demand for a succ…

Cited by 0SourceScholar
2019

Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality

ICASSP 2019accepted

Despite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the trai…

Cited by 0SourceScholar
2018

Towards Conditional Adversarial Training for Predicting Emotions from Speech

ICASSP 2018accepted

Motivated by the encouraging results recently obtained by generative adversarial networks in various image processing tasks, we propose a conditional adversarial training framework to predict dimensional representations of emotion, i. e., arousal and valence, from speech signals. The framework consi…

Cited by 0SourceScholar
2017

Prediction-based learning for continuous emotion recognition in speech

ICASSP 2017accepted

In this paper, a prediction-based learning framework is proposed for a continuous prediction task of emotion recognition from speech, which is one of the key components of affective computing in multimedia. The main goal of this framework is to utmost exploit the individual advantages of different r…

Cited by 0SourceScholar
2017

Reconstruction-error-based learning for continuous emotion recognition in speech

ICASSP 2017accepted

To advance the performance of continuous emotion recognition from speech, we introduce a reconstruction-error-based (RE-based) learning framework with memory-enhanced Recurrent Neural Networks (RNN). In the framework, two successive RNN models are adopted, where the first model is used as an autoenc…

Cited by 0SourceScholar
2016

Enhanced semi-supervised learning for multimodal emotion recognition

ICASSP 2016accepted

Semi-Supervised Learning (SSL) techniques have found many applications where labeled data is scarce and/or expensive to obtain. However, SSL suffers from various inherent limitations that limit its performance in practical applications. A central problem is that the low performance that a classifier…

Cited by 0SourceScholar
2016

Wavelet features for classification of vote snore sounds

ICASSP 2016accepted

Location and form of the upper airway obstruction is essential for a targeted therapy of obstructive sleep apnea (OSA). Utilizing snore sounds (SnS) to reveal the pathological characters of OSA patients has been the subject of scientific research for several decades. Fewer studies exist on the evalu…

Cited by 0SourceScholar