← Search

Xiaojiang Peng

29 accepted papers

2026

Adaptive Action Chunking at Inference-time for Vision-Language-Action Models

CVPR 2026

In Vision-Language-Action (VLA) models, action chunking (i.e., executing a sequence of actions without intermediate replanning) is a key technique to improve robotic manipulation abilities. However, a large chunk size reduces the model's responsiveness to new information, while a small one increases

Cited by 0SourcecodeScholar
2026

EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?

ICLR 2026poster

Descriptive Multimodal Emotion Recognition (DMER) has garnered increasing research attention. Unlike traditional discriminative paradigms that rely on predefined emotion taxonomies, DMER aims to describe human emotional state using free-form natural language, enabling finer-grained and more interpre…

Cited by 0SourcecodeScholar
2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2025

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

ICML 2025oral

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suff…

2025

DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal Retrieval

AAAI 2025technical

With the burst of big data, 2D-3D cross-modal retrieval has received increasing attention, which aims to retrieve relevant data from one modality given the query from the other modality. In this paper, we study an underexplored yet practical problem of semi-supervised 2D-3D cross-modal retrieval, wh…

Cited by 0SourcePDFScholar
2025

DeformAvatar: Point-Based Human Avatar Re-targeting and Rendering

ICASSP 2025accepted

In this paper, we present the DeformAvatar, a novel architecture for human avatar re-targetting and rendering based on point clouds. Given the multiple views of a person, we first build a point-model-paired human representation containing a raw point cloud and an optimal parametric model. Then, we r…

Cited by 0SourceScholar
2025

EA-Vit: Efficient Adaptation for Elastic Vision Transformer

ICCV 2025poster

Vision Transformers (ViTs) have emerged as a foundational model in computer vision, excelling in generalization and adaptation to downstream tasks. However, deploying ViTs to support diverse resource constraints typically requires retraining multiple, size-specific ViTs, which is both time-consuming…

2025

Enhancing 6D Pose Estimation with Cross-modal Fusion Network and Density-peak Keypoint Localization

ICASSP 2025accepted

Current dual-fusion models for 6D pose estimation often lead to increased computational complexity and risk of overfitting with the addition of more networks. To address this, we propose a Cross-modal Fusion Network (CFN), which extracts robust dual-modal features while reducing computation energy a…

Cited by 0SourceScholar
2025

Let Your Features Tell The Differences: Understanding Graph Convolution By Feature Splitting

ICLR 2025poster

Graph Neural Networks (GNNs) have demonstrated strong capabilities in processing structured data. While traditional GNNs typically treat each feature dimension equally important during graph convolution, we raise an important question: **Is the graph convolution operation equally beneficial for each…

2025

MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification

ACL 2025finding

Reasoning is an essential capacity for large language models (LLMs) to address complex tasks, whereas the identification of process errors is vital for improving this ability. Recently, process-level reward models (PRMs) were proposed to provide step-wise rewards that facilitate reinforcement learni…

Cited by 0SourcePDFScholar
2025

Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths

NeurIPS 2025poster

Mamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's cha…

Cited by 0SourceScholar
2025

UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

ICASSP 2025accepted

Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional…

Cited by 0SourceScholar
2025

sEMG-Based Joint Angle Estimation via Hierarchical Spiking Attentional Feature Decomposition Network

RA-L 2025

Surface electromyography (sEMG) has demonstrated significant potential in simultaneous and proportional control (SPC). However, existing algorithms for predicting joint angles based on sEMG often suffer from high inference costs or are limited to specific subjects rather than multi-subject scenarios

Cited by 2SourcecodeScholar
2024

A Challenge Dataset and Effective Models for Conversational Stance Detection

COLING 2024main

Previous stance detection studies typically concentrate on evaluating stances within individual instances, thereby exhibiting limitations in effectively modeling multi-party discussions concerning the same specific topic, as naturally transpire in authentic social media interactions. This constraint…

2024

DSMix: Distortion-Induced Saliency Map Based Pre-training for No-Reference Image Quality Assessment

ECCV 2024poster

"Image quality assessment (IQA) has long been a fundamental challenge in image understanding. In recent years, deep learning-based IQA methods have shown promising performance. However, the lack of large amounts of labeled data in the IQA field has hindered further advancements of these methods. Thi…

2024

Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning

NeurIPS 2024poster

Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, exist…

2024

NuwaDynamics: Discovering and Updating in Causal Spatio-Temporal Modeling

ICLR 2024spotlight

Spatio-temporal (ST) prediction plays a pivotal role in earth sciences, such as meteorological prediction, urban computing. Adequate high-quality data, coupled with deep models capable of inference, are both indispensable and prerequisite for achieving meaningful results. However, the sparsity of da…

Cited by 12SourcePDFScholar
2024

Semi-supervised Knowledge Transfer Across Multi-omic Single-cell Data

NeurIPS 2024poster

Knowledge transfer between multi-omic single-cell data aims to effectively transfer cell types from scRNA-seq data to unannotated scATAC-seq data. Several approaches aim to reduce the heterogeneity of multi-omic data while maintaining the discriminability of cell types with extensive annotated data.…

Cited by 0SourcePDFScholar
2022

An Efficient Training Approach for Very Large Scale Face Recognition

CVPR 2022poster

Face recognition has achieved significant progress in deep learning era due to the ultra-large-scale and welllabeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is indispensable. Th…

Cited by 39PDFcodeScholar
2021

Affordance Transfer Learning for Human-Object Interaction Detection

CVPR 2021poster

Reasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspired by this, we introduce an affordance transfer learning approach to jointly de…

Cited by 139PDFcodeScholar
2021

Detecting Human-Object Interaction via Fabricated Compositional Learning

CVPR 2021poster

Human-Object Interaction (HOI) detection, inferring the relationships between human and objects from images/videos, is a fundamental task for high-level scene understanding. However, HOI detection usually suffers from the open long-tailed nature of interactions with objects, while human has extremel…

Cited by 118PDFcodeScholar
2020

Attention-Driven Dynamic Graph Convolutional Network for Multi-Label Image Recognition

ECCV 2020poster

Recent studies often exploit Graph Convolutional Network (GCN) to model label dependencies to improve recognition accuracy for multi-label image recognition. However, constructing a graph by counting the label co-occurrence possibilities of the training data may degrade model generalizability, espec…

2020

Suppressing Mislabeled Data via Grouping and Self-Attention

ECCV 2020poster

Deep networks achieve excellent results on large-scale clean data but degrade significantly when learning from noisy labels. To suppressing the impact of mislabeled data, this paper proposes a conceptually simple yet efficient training block, termed as Attentive Feature Mixup (AFM), which allows pay…

2020

Suppressing Uncertainties for Large-Scale Facial Expression Recognition

CVPR 2020poster

Annotating a qualitative large-scale facial expression dataset is extremely difficult due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. These uncertainties suspend the progress of large-scale Facial Expression Recognitio…

Cited by 783PDFcodeScholar
2020

Visual Compositional Learning for Human-Object Interaction Detection

ECCV 2020poster

Human-Object interaction (HOI) detection aims to localize and infer relationships between human and objects in an image. It is challenging because an enormous number of possible combinations of objects and verbs types forms a long-tail distribution. We devise a deep Visual Compositional Learning (VC…