← Search

Xiangmin Xu

31 accepted papers

2026

BRIDGECODE: A DUAL SPEECH REPRESENTATION PARADIGM FOR AUTOREGRESSIVE ZERO-SHOT TEXT-TO-SPEECH SYNTHESIS

ICASSP 2026poster

Autoregressive (AR) frameworks have recently achieved remarkable progress in zero-shot text-to-speech (TTS) by leveraging discrete speech tokens and large language model techniques. Despite their success, existing AR-based zero-shot TTS systems face two critical limitations: (i) an inherent speed-qu…

Cited by 0SourcePDFScholar
2026

CLEX: Complementary Label Exchange Learning for Noisy Facial Expression Recognition

CVPR 2026

Facial expression recognition (FER) in the wild is severely hampered by label noise and annotation ambiguity. Existing methods, including sample selection, label ensembling, and consistency regularization, primarily rely on ordinary label supervision and offer limited control over non-target predict

Cited by 0SourceScholar
2026

DAVID: Dual-stage Adaptive Vision-text Integrated Decoupling for Multimodal KV Cache Eviction

AAAI 2026technical

With the rapid development of multimodal large language models (MLLMs), deploying them on low-resource devices remains challenging. Beyond the model size, long multimodal inputs cause substantial memory overhead in the KV cache, making efficient cache management critical. In this paper, we propose D

Cited by 0SourcePDFScholar
2026

HD-PPT: HIERARCHICAL DECODING OF CONTENT- AND PROMPT-PREFERENCE TOKENS FOR INSTRUCTION-BASED TTS

ICASSP 2026oral

Large Language Model (LLM)-based Text-to-Speech (TTS) models have already reached a high degree of naturalness. However, the precision control of TTS inference is still challenging. Although instruction-based Text-to-Speech (Instruct-TTS) models are proposed, these models still lack fine-grained con…

Cited by 0SourcePDFScholar
2026

NeurIPS: Neuro-anatomical Inductive Priors for Sphere-based Brain Decoding

ICML 2026poster

Current fMRI decoders face a performance-fidelity trade-off where efficient ID encoders outperform geometrically-aligned surface-based models. We argue this is an artifact of inefficient surface tokenization and the failure to use anatomy as a predictive signal. We present **NeurIPS**, a framework t…

Cited by 0SourceScholar
2025

Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection

ICCV 2025poster

Open vocabulary Human-Object Interaction (HOI) detection is a challenging task that detects all <human, verb, object> triplets of interest in an image, even those that are not pre-defined in the training set. Existing approaches typically rely on output features generated by large Vision-Language Mo…

2025

CATCH: A Novel Data Synthesis Framework for High Therapy Fidelity and Memory-Driven Planning Chain of Thought in AI Counseling

EMNLP 2025

Recently, advancements in AI counseling based on large language models have shown significant progress. However, existing studies employ a one-time generation approach to synthesize multi-turn dialogue samples, resulting in low therapy fidelity and failing to capture the decision-making rationale be

2025

DecoupledSynth: Enhancing Zero-Shot Text-to-Speech Via Factors Decoupling

ICASSP 2025accepted

Studies of speech representation enhance zero-shot Text-to-Speech by mapping text to intermediate representations before generating speech. However, using representations often struggles to balance linguistic, para-linguistic, and non-linguistic information in speech during the synthesis phase. Addi…

Cited by 0SourceScholar
2025

Drawing Developmental Trajectory from Cortical Surface Reconstruction

ICCV 2025poster

Diffeomorphic-based cortical surface reconstruction typically involves a series of deformation processes to extract the cerebral cortex from brain magnetic resonance images (MRI). While most methods are designed for adult brains using Neural Ordinary Differential Equations (NODE) with fixed step siz…

2025

Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification

CVPR 2025highlight

Text-to-image person re-identification (ReID) aims to retrieve the images of an interested person based on textual descriptions. One main challenge for this task is the high cost in manually annotating large-scale databases, which affects the generalization ability of ReID models. Recent works handl…

2025

PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological Counseling

ACL 2025long

Currently, large language models (LLMs) have made significant progress in the field of psychological counseling. However, existing mental health LLMs overlook a critical issue where they do not consider the fact that different psychological counselors exhibit different personal styles, including lin…

2025

QuantAgents: Towards Multi-agent Financial System via Simulated Trading

EMNLP 2025

In this paper, our objective is to develop a multi-agent financial system that incorporates simulated trading , a technique extensively utilized by financial professionals. While current LLM-based agent models demonstrate competitive performance, they still exhibit significant deviations from real-w

2025

SAKI-RAG: Mitigating Context Fragmentation in Long-Document RAG via Sentence-level Attention Knowledge Integration

EMNLP 2025

Traditional Retrieval-Augmented Generation (RAG) frameworks often segment documents into larger chunks to preserve contextual coherence, inadvertently introducing redundant noise. Recent advanced RAG frameworks have shifted toward finer-grained chunking to improve precision. However, in long-documen

2025

TailorRPA: A Retrieval-Based Framework for Eliciting Personalized and Coherent Role-Playing Agents in General Domain

EMNLP 2025

Recent advancements of general domain oriented Role-playing Agents (RPAs) have enabled the agents to maintain character properties in a wide spectrum of daily tasks beyond mere scenario based chit-chatting. Nonetheless, current works lacks consideration of replicating internal properties of characte

Cited by 0SourcePDFScholar
2025

TreeRAG: Unleashing the Power of Hierarchical Storage for Enhanced Knowledge Retrieval in Long Documents

ACL 2025finding

When confronting long document information retrieval for Query-Focused Summarization(QFS), Traditional Retrieval-Augmented Generation(RAG) frameworks struggle to retrieve all relevant knowledge points, and the chunking and retrieve strategies of existing frameworks may disrupt the connections betwee…

Cited by 0SourcePDFScholar
2025

Understanding Dynamic Human-Robot Proxemics in the Case of Four-Legged Canine-Inspired Robots

ICRA 2025

The integration of humanoid and animal-shaped robots into specialized domains, such as healthcare, multiterrain operations, and psychotherapy, necessitates a deep understanding of proxemics-the study of spatial behavior that governs effective human-robot interactions. Unlike traditional robots in ma

Cited by 11SourceScholar
2024

Clinical Scores Prediction and Medication Adjustment for Course of Parkinson's Disease

ICASSP 2024accepted

Parkinson's Disease (PD) is the second most prevalent neurodegenerative disorder worldwide, characterized by progressive motor and non-motor symptoms. Unfortunately, there are no definitive PD modifying therapies, so accurate course prediction in advance and appropriate medical adjustment are essent…

Cited by 0SourceScholar
2024

Exploring 3D Human Pose Estimation and Forecasting from the Robot’s Perspective: The HARPER Dataset

IROS 2024poster

We introduce HARPER, a novel dataset for 3D body pose estimation and forecasting in dyadic interactions between users and Spot, the quadruped robot manufactured by Boston Dynamics. The key-novelty of HARPER is its focus on the robot’s perspective, i.e., on the data captured by the robot’s sensors. T…

Cited by 3SourceScholar
2024

Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-On

CVPR 2024poster

Image-based virtual try-on is an increasingly important task for online shopping. It aims to synthesize images of a specific person wearing a specified garment. Diffusion model-based approaches have recently become popular as they are excellent at image synthesis tasks. However these approaches usua…

2023

DST: Deformable Speech Transformer for Emotion Recognition

ICASSP 2023accepted

Enabled by multi-head self-attention, Transformer has exhibited remarkable results in speech emotion recognition (SER). Compared to the original full attention mechanism, window-based attention is more effective in learning fine-grained features while greatly reducing model redundancy. However, emot…

Cited by 0SourceScholar
2023

DWFormer: Dynamic Window Transformer for Speech Emotion Recognition

ICASSP 2023accepted

Speech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreover, the temporal scales of important information may vary over a large range within and across speech segments. Although…

Cited by 0SourceScholar
2023

MGAT: Multi-Granularity Attention Based Transformers for Multi-Modal Emotion Recognition

ICASSP 2023accepted

Multi-modal emotion recognition is crucial for human-computer interaction. Many existing algorithms attempt to achieve multi-modal interactions through a cross-attention mechanism. Due to the problems of noise introduction and heavy computation in the original attention mechanism, window attention h…

Cited by 0SourceScholar
2023

SoulChat: Improving LLMs' Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations

EMNLP 2023short findings

Large language models (LLMs) have been widely applied in various fields due to their excellent capability for memorizing knowledge and chain of thought (CoT). When these language models are applied in the field of psychological counseling, they often rush to provide universal advice. However, when u…

Cited by 0SourcecodeScholar
2023

Speaker-Aware Hierarchical Transformer For Personality Recognition In Multiparty Dialogues

ICASSP 2023accepted

Personality recognition is one of the core technologies in human-machine interaction, which has received increasing attention. Previous works mainly focus on essays or monologues, while personality traits reveal more in the interactions with others. Due to the lack of appropriate datasets, a few app…

Cited by 0SourceScholar
2023

Superpoint Transformer for 3D Scene Instance Segmentation

AAAI 2023technical

Most existing methods realize 3D instance segmentation by extending those models used for 3D object detection or 3D semantic segmentation. However, these non-straightforward methods suffer from two drawbacks: 1) Imprecise bounding boxes or unsatisfactory semantic predictions limit the performance of…

2022

CS-GResNet: A Simple and Highly Efficient Network for Facial Expression Recognition

ICASSP 2022accepted

Facial expression recognition (FER) has recently attracted attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources and memory consumption. Hence, achieving promising performance while maintaining the efficiency of mo…

Cited by 0SourceScholar
2022

Key-Sparse Transformer for Multimodal Speech Emotion Recognition

ICASSP 2022accepted

Speech emotion recognition is a challenging research topic that plays a critical role in human-computer interaction. Multimodal inputs further improve the performance as more emotional information is used. However, existing studies learn all the information in the sample while only a small portion o…

Cited by 0SourceScholar
2021

LSSED: A Large-Scale Dataset and Benchmark for Speech Emotion Recognition

ICASSP 2021accepted

Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED, a challenging large-scale english speech emotion dataset, w…

Cited by 0SourceScholar
2018

Recurrent Neural Networks for Automatic Replay Spoofing Attack Detection

ICASSP 2018accepted

In order to enhance the security of automatic speaker verification (ASV) systems, automatic spoofing attack detection, which discriminates the fake audio recordings from genuine human speech, has gain much attention recently. Among various ways of spoofing attacks, replay attacks are one of the most…

Cited by 0SourceScholar