← Search

Pengxiang Ding

29 accepted papers

2026

CUBic: Coordinated Unified Bimanual Perception and Control Framework

CVPR 2026

Recent advances in visuomotor policy learning have enabled robots to perform control directly from visual inputs. Yet, extending such end-to-end learning from single-arm to bimanual manipulation remains challenging due to the need for both independent perception and coordinated interaction between a

Cited by 0SourceScholar
2026

Dyn-VPP: Video Prediction Policy Optimization for Improved Visual Dynamics

ICML 2026poster

Video action models are a promising foundation for Vision–Language–Action (VLA) because they can learn rich visual dynamics directly from video. However, likelihood-oriented training of diffusion predictors emphasizes globally plausible futures and does not guarantee precision-critical visual dynami…

Cited by 0SourceScholar
2026

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering from temporal myopia that degrades long-horizon coherence. In

Cited by 0SourcecodeScholar
2026

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

AAAI 2026technical

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention to target regions. Instead, visual attention is always dispe

Cited by 0SourcePDFScholar
2026

Rethinking Target Label Conditioning in Adversarial Attacks: A 2D Tensor-Guided Generative Approach

AAAI 2026technical

Compared to single-target adversarial attacks, multi-target attacks have garnered significant attention due to their ability to generate adversarial images for multiple target classes simultaneously. However, existing generative approaches for multi-target attacks primarily encode target labels into

Cited by 0SourcePDFScholar
2026

Rethinking the Practicality of Vision-Language-Action Model: A Comprehensive Benchmark and an Improved Baseline

ICRA 2026poster

Vision-Language-Action (VLA) models have emerged as a generalist robotic agent. However, existing VLAs are hindered by excessive parameter scales, prohibitive pre-training requirements, and limited applicability to diverse embodiments. To improve the practicality of VLAs, we propose a comprehensive …

2026

Robust Online Residual Refinement Via Koopman-Guided Dynamics Modeling

ICRA 2026poster

Imitation learning (IL) enables efficient skill acquisition from demonstrations but often struggles with long-horizon tasks and high-precision control due to compounding errors. Residual policy learning offers a promising, model-agnostic solution by refining a base policy through closed-loop correct…

2026

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

ICLR 2026poster

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data, which lack accurate spatial awareness and hinder their abili…

Cited by 0SourcecodeScholar
2026

TrajBooster: Boosting Humanoid Whole-Body Manipulation Via Trajectory-Centric Learning

ICRA 2026poster

Recent Vision-Language-Action (VLA) models show potential to generalize across embodiments but struggle to quickly align with a new robot’s action space when high-quality demonstrations are scarce, especially for bipedal humanoids. We present TrajBooster, a cross-embodiment framework that leverages …

2026

Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Diffusion Diffusion Process

ICLR 2026poster

Vision-language-action (VLA) models aim to understand natural language instructions and visual observations and execute corresponding actions as an embodied agent. Recent advancements have integrated future images into the understanding-action loop, enabling foresight-driven policies that reduce abs…

Cited by 0SourcecodeScholar
2026

VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

AAAI 2026technical

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how

Cited by 0SourcePDFScholar
2025

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction

ICCV 2025accepted

In robotic visuomotor policy learning, diffusion-based models have achieved significant success in improving the accuracy of action trajectory generation compared to traditional autoregressive models. However, they suffer from inefficiency due to multiple denoising steps and limited flexibility from…

Cited by 0SourcePDFScholar
2025

Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

AAAI 2025technical

In recent years, applying multi-modal large language models (MLLMs) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, MLLMs comprise the well-known Transformer network, which has a less efficient quadratic computation complexity. In this s…

2025

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

ICLR 2025poster

With the rapid development of embodied artificial intelligence, significant progress has been made in vision-language-action (VLA) models for general robot decision-making. However, the majority of existing VLAs fail to account for the inevitable external perturbations encountered during deployment.…

Cited by 2SourcePDFScholar
2025

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

CoRL 2025poster

Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing VLA frameworks primarily address short-horizon tasks, and their effectiveness on long-horizon, multi-step robotic manipu…

Cited by 0SourceScholar
2025

MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models

ICRA 2025

Developing versatile quadruped robots that can smoothly perform various actions and tasks in real-world environments remains a significant challenge. This paper introduces a novel vision-language-action (VLA) model, mixture of robotic experts (MoRE), for quadruped robots that aim to introduce reinfo

Cited by 24SourceScholar
2025

PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding

IROS 2025

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control. However, action chunking linearly scales up action dimensions in

Cited by 60SourceScholar
2025

Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning

ICRA 2025

This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the l

Cited by 1SourcecodeScholar
2025

ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning

ICML 2025poster

Vision-Language-Action (VLA) models have shown great potential in general robotic decision-making tasks via imitation learning. However, the variable quality of training data often constrains the performance of these models. On the other hand, offline Reinforcement Learning (RL) excels at learning r…

Cited by 0SourcePDFScholar
2025

Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation

ICML 2025poster

Behavior Cloning (BC) is a widely adopted visual imitation learning method in robot manipulation. Current BC approaches often enhance generalization by leveraging large datasets and incorporating additional visual and textual modalities to capture more diverse information. However, these methods ove…

Cited by 0SourcePDFScholar
2025

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

NeurIPS 2025poster

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clouds or depth, either require specialized sensors or fail to effectively exploit d…

Cited by 0SourcecodeScholar
2025

Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport

ICML 2025poster

Diffusion policies have shown promise in learning complex behaviors from demonstrations, particularly for tasks requiring precise control and long-term planning. However, they face challenges in robustness when encountering distribution shifts. This paper explores improving diffusion-based imitation…

2025

VLAS: Vision-Language-Action Model with Speech Instructions for Customized Robot Manipulation

ICLR 2025poster

Vision-language-action models (VLAs) have recently become highly prevalent in robot manipulation due to its end-to-end architecture and impressive performance. However, current VLAs are limited to processing human instructions in textual form, neglecting the more natural speech modality for human in…

2024

Expressive Forecasting of 3D Whole-Body Human Motions

AAAI 2024technical

Human motion forecasting, with the goal of estimating future human behavior over a period of time, is a fundamental task in many real-world applications. However, existing works typically concentrate on foretelling the major joints of the human body without considering the delicate movements of the…

2024

GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot

IROS 2024poster

Multi-task robot learning holds significant importance in tackling diverse and complex scenarios. However, current approaches are hindered by performance issues and difficulties in collecting training datasets. In this paper, we propose GeRM (Generalist Robotic Model). We utilize offline reinforceme…

Cited by 13SourcecodeScholar
2024

PiTe: Pixel-Temporal Alignment for Large Video-Language Model

ECCV 2024oral

"Fueled by the Large Language Models (LLMs) wave, Large Visual-Language Models (LVLMs) have emerged as a pivotal advancement, bridging the gap between image and text. However, video making it challenging for LVLMs to perform adequately due to the complexity of the relationship between language and s…

2024

QUAR-VLA: Vision-Language-Action Model for Quadruped Robots

ECCV 2024poster

"The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional quadruped robot learning typically handles language interaction and visual autonomous perception separately, which, while simplifying system design, also limits the sy…

Cited by 19SourcePDFScholar
2024

RL2AC: Reinforcement Learning-based Rapid Online Adaptive Control for Legged Robot Robust Locomotion

RSS 2024poster

Dynamic fast adaptation is one of the basic capabilities that enables the animals to timely and properly adjust its locomotion reacting to the unpredictable changes. Such capability is also essential for the quadruped robot, when working in the unforseen environment. While reinforcement learning (RL…

Cited by 6SourcePDFScholar