← Search

Xiao-Ping Zhang

48 accepted papers

2026

AVR: Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization

ICRA 2026poster

Robotic manipulation in complex scenes demands precise perception of task-relevant details, yet fixed or suboptimal viewpoints often impair fine-grained perception and induce occlusions, constraining imitation-learned policies. We present AVR (Active Vision-driven Robotics), a bimanual teleoperation…

2026

LEND A HAND: SEMI TRAINING-FREE CUED SPEECH RECOGNITION VIA MLLM-DRIVEN HAND MODELING FOR BARRIER-FREE COMMUNICATION

ICASSP 2026poster

Cued Speech (CS) is an innovative visual communication system that integrates lip-reading with hand coding, designed to enhance effective communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) refers to the AI-driven process of automatically recognizing hand gestures…

Cited by 0SourcePDFScholar
2026

LearniBridge: Learnable Calibration of Feature Caching for Diffusion Models Acceleration

ICML 2026poster

Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs. Feature caching accelerates inference by reusing intermediate representations. Existing methods rely on historical features for implementation simplicity, yet…

Cited by 0SourceScholar
2026

MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs

ICLR 2026poster

Developing Large Language Models (LLMs) to cooperate and compete effectively within multi-agent systems (MASs) is a critical step towards more advanced intelligence. While reinforcement learning (RL) has proven effective for enhancing reasoning in single-agent tasks, its extension to multi-turn, mul…

Cited by 13SourcecodeScholar
2026

SCOPE: Skeleton Graph-Based Computation-Efficient Framework for Autonomous UAV Exploration

RA-L 2026

Autonomous exploration in unknown environments is key for mobile robots, helping them perceive, map, and make decisions in complex areas. However, current methods often rely on frequent global optimization, suffering from high computational latency and trajectory oscillation, especially on resource-

Cited by 0SourceScholar
2025

ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models

CVPR 2025poster

Large Vision Language Models (LVLMs) have achieved significant success across multi-modal tasks. However, the computational cost of processing long visual tokens can be prohibitively expensive on resource-limited devices. Previous methods have identified redundancy in visual tokens within the Large…

Cited by 8SourcePDFScholar
2025

CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression Segmentation

ICCV 2025poster

The newly proposed Generalized Referring Expression Segmentation (GRES) amplifies the formulation of classic RES by involving complex multiple/non-target scenarios. Recent approaches address GRES by directly extending the well-adopted RES frameworks with object-existence identification. However, the…

Cited by 0SourcePDFScholar
2025

Diff2I2P: Differentiable Image-to-Point Cloud Registration with Diffusion Prior

ICCV 2025poster

Learning cross-modal correspondences is essential for image-to-point cloud (I2P) registration. Existing methods achieve this mostly by utilizing metric learning to enforce feature alignment across modalities, disregarding the inherent modality gap between image and point data. Consequently, this par…

2025

Distributional Decision Transformer: Risk-Sensitive Offline RL via Quantile-Based Critics and Stochastic Return

IROS 2025

Offline reinforcement learning faces a critical challenge in synthesizing high-reward trajectories from suboptimal datasets while robustly handling the stochasticity inherent in real-world decision-making. While combination of return-conditioned sequence models, such as Decision Transformers (DT), a

Cited by 0SourceScholar
2025

Exo-ViHa: A Cross-Platform Exoskeleton System with Visual and Haptic Feedback for Efficient Dexterous Skill Learning

IROS 2025

Imitation learning has emerged as a powerful paradigm for robot skills learning. However, traditional data collection systems for dexterous manipulation face challenges, including a lack of balance between acquisition efficiency, consistency, and accuracy. To address these issues, we introduce Exo-V

Cited by 3SourcecodeScholar
2025

GLST-GCN: Global-Local Spatio-Temporal Graph Convolutional Network for Skeleton-based Hand Motion Prediction

ICASSP 2025accepted

This paper introduces a new task: skeleton-based hand motion sequence prediction, which can be applied to VR/AR systems and human-computer interaction systems to enhance user experience. To tackle this task, we performed a comprehensive analysis of hand movement patterns and propose a Global-Local S…

Cited by 0SourceScholar
2025

GauUpdate: New Object Insertion in 3D Gaussian Fields with Consistent Global Illumination

ICCV 2025poster

3D Gaussian Splatting (3DGS) is a prevailing technique to reconstruct large-scale 3D scenes from multiview images for novel view synthesis, like a room, a block, and even a city. Such large-scale scenes are not static with changes constantly happening in these scenes, like a new building being built…

Cited by 0SourcePDFScholar
2025

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Task

ACL 2025finding

In this paper, we present HumanEval Pro and MBPP Pro, a series of benchmarks to evaluate LLMs on self-invoking code generation task. This task involves providing LLMs with a base problem alongside a related, more complex problem. The models must solve the base problem and leverage its solution to ad…

2025

Mean-Field Aided QMIX: A Scalable and Flexible Q-Learning Approach for Large-Scale Agent Groups

ICASSP 2025accepted

Value decomposition methods are effective for multi-agent reinforcement learning (MARL), with QMIX being one of the most advanced. However, it struggles with scalability and flexibility in large-scale agent systems. The introduction of mean-field theory into MARL provides a potential solution to bot…

Cited by 0SourceScholar
2025

Multi-View 3D Human Pose Estimation with Weakly Synchronized Images

AAAI 2025technical

Multi-view 3D human pose estimation (MHPE) is an important research task in computer vision. To maintain consistency during the data collection, hardware synchronization devices are commonly used to connect cameras, ensuring that images from different views are captured simultaneously. However, sync…

Cited by 0SourcePDFScholar
2025

PID-controlled Langevin Dynamics for Faster Sampling on Generative Models

NeurIPS 2025poster

Langevin dynamics sampling suffers from extremely low generation speed, fundamentally limited by numerous fine-grained iterations to converge to the target distribution. We introduce PID-controlled Langevin Dynamics (PIDLD), a novel sampling acceleration algorithm that reinterprets the sampling proc…

Cited by 0SourcecodeScholar
2025

Reinforced Domain Selection for Continuous Domain Adaptation

ICASSP 2025accepted

Continuous Domain Adaptation (CDA) effectively bridges significant domain shifts by progressively adapting from the source domain through intermediate domains to the target domain. However, selecting intermediate domains without explicit metadata remains a substantial challenge that has not been ext…

Cited by 0SourceScholar
2025

SynTSBench: Rethinking Temporal Pattern Learning in Deep Learning Models for Time Series

NeurIPS 2025poster

Recent advances in deep learning have driven rapid progress in time series forecasting, yet many state-of-the-art models continue to struggle with robust performance in real-world applications, even when they achieve strong results on standard benchmark datasets. This persistent gap can be attribute…

Cited by 0SourceScholar
2025

Transfer Risk Map: Mitigating Pixel-level Negative Transfer in Medical Segmentation

ICASSP 2025accepted

How to mitigate negative transfer in transfer learning is a long-standing and challenging issue, especially in the application of medical image segmentation. Existing methods for reducing negative transfer focuses on classification or regression tasks, ignoring the non-uniform negative transfer risk…

Cited by 0SourceScholar
2024

Dual-modal Tactile E-skin: Enabling Bidirectional Human-Robot Interaction via Integrated Tactile Perception and Feedback

ICRA 2024poster

To foster an immersive and natural human-robot interaction (HRI), the implementation of tactile perception and feedback becomes imperative, effectively bridging the conventional sensory gap. In this paper, we propose a dual-modal electronic skin (e-skin) that integrates magnetic tactile sensing and…

Cited by 2SourceScholar
2024

Language-Free Compositional Action Generation via Decoupling Refinement

ICASSP 2024accepted

Composing simple actions into complex actions is crucial yet challenging. Existing methods largely rely on language annotations to discern composable latent semantics, which is costly and labor-intensive. In this study, we introduce a novel framework to generate compositional actions without languag…

Cited by 0SourceScholar
2024

Leveraging Noisy Labels of Nearest Neighbors for Label Correction and Sample Selection

ICASSP 2024accepted

Dealing with noisy labels (LNL) emerges as a critical challenge when applying deep learning (DL) in practical settings. Previous methodologies primarily concentrated on harnessing model predictions to mitigate the impact of noisy labels. Nevertheless, their efficacy is strongly contingent on the acc…

Cited by 0SourceScholar
2024

Multi-scale Consistency for Robust 3D Registration via Hierarchical Sinkhorn Tree

NeurIPS 2024poster

We study the problem of retrieving accurate correspondence through multi-scale consistency (MSC) for robust point cloud registration. Existing works in a coarse-to-fine manner either suffer from severe noisy correspondences caused by unreliable coarse matching or struggle to form outlier-free coarse…

Cited by 0SourcePDFScholar
2024

Optimizing Trading Strategies in Quantitative Markets Using Multi-Agent Reinforcement Learning

ICASSP 2024accepted

Quantitative markets are characterized by swift dynamics and abundant uncertainties, making the pursuit of profit-driven stock trading actions inherently challenging. Within this context, Reinforcement Learning (RL) — which operates on a reward-centric mechanism for optimal control — has surfaced as…

Cited by 0SourceScholar
2024

SATac: A Thermoluminescence Enabled Tactile Sensor for Concurrent Perception of Temperature, Pressure, and Shear

ICRA 2024poster

Most vision-based tactile sensors use elastomer deformation to infer tactile information, which can not sense some modalities, like temperature. As an important part of human tactile perception, temperature sensing can help robots better interact with the environment. In this work, we propose a nove…

Cited by 1SourceScholar
2024

Social Physics Informed Diffusion Model for Crowd Simulation

AAAI 2024technical

Crowd simulation holds crucial applications in various domains, such as urban planning, architectural design, and traffic arrangement. In recent years, physics-informed machine learning methods have achieved state-of-the-art performance in crowd simulation but fail to model the heterogeneity and mul…

2024

Unified Probability Distributions of Generalized Composite Fading with Inverse-Type Distributions of Large-Scale Shadowing/Fluctuations

ICASSP 2024accepted

Based on novel inverse-type PDF formulae for large-scale shadowing/fluctuations, we derive novel unified probability density functions (PDFs) and moment generating function (MGF) formulae that characterize wide ranging of generalized composite fading distributions in radio frequency and free-space o…

Cited by 0SourceScholar
2023

Autonomous Swarm Robot Coordination via Mean-Field Control Embedding Multi-Agent Reinforcement Learning

IROS 2023poster

The learning approaches of designing a controller to guide the collective behavior of swarm robots have gained significant attention in recent years. However, the scalability of swarm robots and their inherent stochasticity complicate the control problem due to increasing complexity, unpredictabilit…

Cited by 4SourceScholar
2023

Iterative Water-Filling Power and Subcarrier Allocation for Multicarrier NOMA Downlink

ICASSP 2023accepted

Novel closed form formulae of iterative optimal power control and allocation, and criterion for optimal subcarrier allocation are derived for downlink of multicarrier non-orthogonal multiple access (MC-NOMA) systems. For the first time, we present closed form water-filling formulae that quantify exa…

Cited by 0SourceScholar
2023

STEV: Stretchable Triboelectric E-skin enabled Proprioceptive Vibration Sensing for Soft Robot

ICRA 2023poster

Vibration perception is essential for robotic sensing and dynamic control. Nevertheless, due to the rigorous demand for sensor conformability and stretchability, enabling soft robots with proprioceptive vibration sensing remains challenging. This paper proposes a novel liquid metal-based stretchable…

Cited by 6SourceScholar
2023

TEFISTA-NET: GTD Parameter Estimation of Low-Frequency Ultra- Wideband Radar via Model-Based Deep Learning

ICASSP 2023accepted

The geometrical theory of diffraction (GTD) has been widely investigated to describe the target scattering behaviors with the low-frequency ultra-wideband (LFW) radar. In this paper, we propose a new model-based deep learning method for GTD parameter estimation. The proposed method is designed by un…

Cited by 0SourceScholar
2023

Tem-Adapter: Adapting Image-Text Pretraining for Video Question Answer

ICCV 2023poster

Video-language pre-trained models have shown remarkable success in guiding video question-answering (VideoQA) tasks. However, due to the length of video sequences, training large-scale video-based models incurs considerably higher costs than training image-based ones. This motivates us to leverage t…

Cited by 17PDFcodeScholar
2023

Unobtrusive Respiratory Monitoring System for Intensive Care

ICASSP 2023accepted

The video-based non-contact respiration detection technology can be used in many application scenarios to unobtrusively and ubiquitously monitor the physical state of living beings, and various researchers are currently working on this technology. The optical flow method in tandem with crossover poi…

Cited by 0SourceScholar
2023

Visuotactile Sensor Enabled Pneumatic Device Towards Compliant Oropharyngeal Swab Sampling

IROS 2023poster

Manual oropharyngeal (OP) swab sampling is an intensive and risky task. In this article, a novel OP swab sampling device of low cost and high compliance is designed by combining the visuotactile sensor and the pneumatic actuator-based gripper. Here, a concave visuotactile sensor called CoTac is firs…

Cited by 3SourceScholar
2022

Planar Magnetic Actuation for Soft and Rigid Robots Using a Scalable Electromagnet Array

RA-L 2022

Magnetic actuation system manipulates micro soft or rigid robots by a controllable magnetic field to move them freely in the narrow or enclosed space, which has demonstrated its huge potential in medical interventional surgery and drug delivery. However, the limited working space of paired or area-c

Cited by 16SourceScholar
2021

An Attention-Seq2Seq Model Based on CRNN Encoding for Automatic Labanotation Generation from Motion Capture Data

ICASSP 2021accepted

Labanotation is an important notation system widely used for recording dances. Numerous methods have been proposed for automatic Labanotation generation from motion capture data. Recently, the sequence-to-sequence (seq2seq) model is proposed. However, the encoder of the model only encodes the tempor…

Cited by 0SourceScholar
2021

Boosting Video Representation Learning With Multi-Faceted Integration

CVPR 2021poster

Video content is multifaceted, consisting of objects, scenes, interactions or actions. The existing datasets mostly label only one of the facets for model training, resulting in the video representation that biases to only one facet depending on the training dataset. There is no study yet on how to…

Cited by 13PDFScholar
2021

Direction-aware Feature-level Frequency Decomposition for Single Image Deraining

IJCAI 2021poster

We present a novel direction-aware feature-level frequency decomposition network for single image deraining. Compared with existing solutions, the proposed network has three compelling characteristics. First, unlike previous algorithms, we propose to perform frequency decomposition at feature-level…

Cited by 3SourcePDFScholar
2021

Explainable Person Re-Identification With Attribute-Guided Metric Distillation

ICCV 2021poster

Despite the great progress of person re-identification (ReID) with the adoption of Convolutional Neural Networks, current ReID models are opaque and only outputs a scalar distance between two persons. There are few methods providing users semantically understandable explanations for why two persons…

Cited by 57PDFcodeScholar
2021

Exploiting Relationship for Complex-scene Image Generation

AAAI 2021technical

The significant progress on Generative Adversarial Networks (GANs) has facilitated realistic single-object image generation based on language input. However, complex-scene generation (with various interactions among multiple objects) still suffers from messy layouts and object distortions, due to di…

2021

Identification of Deep Breath While Moving Forward Based on Multiple Body Regions and Graph Signal Analysis

ICASSP 2021accepted

This paper presents an unobtrusive solution that can automatically identify deep breath when a person is walking past the global depth camera. Existing non-contact breath assessments achieve satisfactory results under restricted conditions when human body stays relatively still. When someone moves f…

Cited by 0SourceScholar
2021

Optimal TOA Localization for Moving Sensor in Asymmetric Network

ICASSP 2021accepted

In a localization system based-on asymmetric network, only one of the anchor nodes (ANs) transmits signal. A sensor node (SN) receives it and then transmits signal that is received by all ANs to form time-of-arrival (TOA) measurements. SN localization is achieved based-on these TOA measurements alon…

Cited by 0SourceScholar
2021

Perceptual Quality Assessment for Recognizing True and Pseudo 4k Content

ICASSP 2021accepted

To meet the imperative demand for monitoring the quality of Ultra High-Definition (UHD) content in multimedia industries, we propose an efficient no-reference (NR) image quality assessment (IQA) metric to distinguish original and pseudo 4K contents and measure the quality of their quality in this pa…

Cited by 0SourceScholar
2016

A Weighted Variational Model for Simultaneous Reflectance and Illumination Estimation

CVPR 2016poster

We propose a weighted variational model to estimate both the reflectance and the illumination from an observed image. We show that, though it is widely adopted for ease of modeling, the log-transformed image for this task is not ideal. Based on the previous investigation of the logarithmic transform…

Cited by 1183PDFScholar