← Search

ZhiPeng Wang

31 accepted papers

2026

$\textit{S}$-SPPO: Semantic-Calibrated Self-Play Preference Optimization

ICML 2026poster

Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley-Terry instantiation of DPO is limited in modeling common departures from transitivity in human preferences. To address this, recent work has introd…

Cited by 0SourceScholar
2026

Autonomous UAV–Quadruped Docking in Complex Terrains Via Active Posture Alignment and Constraint-Aware Control

ICRA 2026poster

Autonomous docking between Unmanned Aerial Vehicles (UAVs) and ground robots is essential for heterogeneous systems, yet most existing approaches target wheeled platforms whose limited mobility constrains exploration in complex terrains. Quadruped robots offer superior adaptability but undergo frequ…

2026

Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction

ICRA 2026poster

Cutting-edge robot learning techniques including foundation models and imitation learning from humans all pose huge demands on large-scale and high-quality datasets which constitute one of the bottleneck in the general intelligent robot fields. This paper presents the Kaiwu multimodal dataset to add…

2026

LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding

CVPR 2026

Autoregressive models (ARMs) have long dominated the landscape of biomedical vision-language models (VLMs). Recently, masked diffusion models such as LLaDA have emerged as promising alternatives, yet their application in the biomedical domain remains largely underexplored. To bridge this gap, we int

Cited by 0SourcecodeScholar
2026

LungNoduleAgent: A Collaborative Multi-Agent System for Precision Diagnosis of Lung Nodules

AAAI 2026technical

Diagnosing lung cancer typically involves physicians identifying lung nodules in Computed tomography (CT) scans and generating diagnostic reports based on their morphological features and medical expertise. Although advancements have been made in using multimodal large language models for analyzing

Cited by 0SourcePDFScholar
2026

Morphogenetic Assembly and Adaptive Control for Heterogeneous Modular Robots

ICRA 2026poster

This paper presents a closed-loop automation framework for heterogeneous modular robots, encompassing the entire pipeline from morphological construction to adaptive control. Within this framework, a mobile manipulator manipulates heterogeneous functional modules—including structural, joint, and whe…

2026

Offline-Trained GAN-Augmented Highly Adaptive Control with Multi-DoF Fusion for Pneumatic Soft Surgical Robots (I)

ICRA 2026poster

Pneumatic soft robots are well-suited for minimally invasive surgery owing to their compliance and safe interaction with tissues. However, achieving highly adaptive control is difficult owing to modeling inaccuracies, inter-chamber coupling, and disturbances from surgical instruments. Non-learning a…

Cited by 0Scholar
2026

Orthogonal Ray Projection: A Tangent-Space Visual Measurement Model for Robust Visual-Inertial Odometry

RA-L 2026

The reprojection error in Visual-Inertial Odometry (VIO) suffers from high nonlinearity due to perspective division, which degrades estimator consistency and robustness, particularly under large depth uncertainty. To address this, we propose a novel visual measurement model, the Orthogonal Ray Proje

Cited by 0SourceScholar
2026

PaiP: An Operational Aware Interactive Planner for Unknown Cabinet Environments

ICRA 2026poster

Box/cabinet scenarios pose with stacked objects significant challenges for robotic motion due to visual occlusions and constrained free space. Traditional collision-free trajectory planning methods often fail when no collision-free paths exist, and may even lead to catastrophic collisions caused by …

2026

Predicting Tactile Sensory Outcome of Physical Human-Robot Interaction Through Embodied Learning Strategy

RA-L 2026

Estimation of robotic dynamic states in physical human-robot interaction (pHRI) is crucial for robots to handle real-world uncertainties. Since wearable tactile sensor arrays have been increasingly applied in robots, the tactile sensory outcomes (TSOs) during pHRIs would provide practical and ideal

Cited by 0SourceScholar
2026

Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction

ICLR 2026poster

Reasoning language models such as DeepSeek-R1 produce long chain-of-thought traces during inference time which make them costly to deploy at scale. We show that using compression techniques such as neural network pruning produces greater performance loss than in typical language modeling tasks, and…

Cited by 0SourcecodeScholar
2026

VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation

ICRA 2026poster

In the context of imitation learning, visuomotor-based diffusion policy learning is one of the main directions in robotic manipulation. Most of these approaches rely on point clouds as observation inputs and construct scene representations through point clouds feature learning, which enables them to…

2026

When Birds Meet Fish: Vision-Force Fusion for Autonomous Underwater Docking in Cross-Domain Avian-Aquatic Collaboration

ICRA 2026poster

Unmanned aerial–aquatic vehicles (UAAVs) provide cross-domain adaptability and broad visions, while autonomous underwater vehicles (AUVs) support long-duration operations. This work integrates the two by developing a rapid underwater docking and releasing system. An autonomous clamping mechanism is …

Cited by 0Scholar
2025

From End-to-end to Step-by-step: Learning to Abstract via Abductive Reinforcement Learning

IJCAI 2025

Abstraction is a critical technique in general problem-solving, allowing complex tasks to be decomposed into smaller, manageable sub-tasks. While traditional symbolic planning relies on predefined primitive symbols to construct structured abstractions, its reliance on formal representations limits a

2025

Human-in-the-Loop Optimization for Knee Exoskeleton Flexion Assistance

RA-L 2025

Human-in-the-loop optimization (HILO) has been used to identify subject-specific assistive strategies and improve the performance of wearable exoskeletons. However, there is still a gap in research on HILO regarding knee exoskeleton flexion assistance. We present a HILO methodology that optimizes th

Cited by 5SourceScholar
2025

Learning Efficient Robotic Garment Manipulation with Standardization

ICML 2025poster

Garment manipulation is a significant challenge for robots due to the complex dynamics and potential self-occlusion of garments. Most existing methods of efficient garment unfolding overlook the crucial role of standardization of flattened garments, which could significantly simplify downstream task…

2025

MUCD: Unsupervised Point Cloud Change Detection via Masked Consistency

AAAI 2025technical

3D Change Detection (3DCD) has gradually become another research hotspot after image change detection. Recent works focus on using artificial labels for supervised or weakly-supervised training of siamese networks to segment changed points. However, labeling every points of multi-temporal point clou…

Cited by 0SourcePDFScholar
2025

NeuTRL: Neural Trust-Guided Reinforcement Learning for Human-Robot Collaboration

RA-L 2025

Reinforcement Learning from Human Feedback (RLHF) enables robots to learn cooperative strategies aligned with human expectations by incorporating feedback into the learning process. However, existing RLHF methods rely on explicit query-based feedback, which is limited for complex, long-horizon tasks

Cited by 6SourceScholar
2025

Rotation Invariant Spatial Networks for Single-View Point Cloud Classification

IJCAI 2025

Point cloud classification is critical for three-dimensional scene understanding. However, in real-world scenarios, depth cameras often capture partial, single-view point clouds of objects with different poses, making their accurate classification a challenge. In this paper, we propose a novel point

2025

Sensing Differently: Unifying Vision, Language, Posture and Tactile in Robotic Perception

IROS 2025

Multi-modal fusion perception enhances robotic performance in complex tasks by providing more comprehensive information than single modality. While tactile and proprioceptive sensing are effective for direct contact tasks like grasping, current research mainly focuses on vision-language fusion, negl

Cited by 0SourceScholar
2025

Uni-Zipper: A Multi-modal Perception Framework of Deformable Objects with Unpaired Data

IROS 2025

Multi-modal perception plays a crucial role in preventing deformation and damage during the robotic manipulation of deformable objects. However, integrating new heterogeneous modalities into existing robotic perception frameworks remains a significant challenge, primarily due to the need for massive

Cited by 0SourceScholar
2024

Learning Cross Dimension Scene Representation for Interactive Navigation Agents in Obstacle-Cluttered Environments

RA-L 2024

Embodied visual navigation has witnessed significant advancements. However, most studies commonly assume that environments are static and contain at least one collision-free path. In human environments, agents frequently encounter challenges when navigating through scenes with disarranged objects. I

Cited by 2SourceScholar
2024

Neighbor Similarity and Multimodal Alignment based Product Recommendation Study

UAI 2024poster

Existing multimodal recommendation research still faces some challenges, such as not being able to fully mine the implicit relevance information of neighbor nodes, and the unreasonable weight allocation to imbalanced nodes. To address the aforementioned challenges, this paper introduces a new multim…

Cited by 0SourcePDFScholar
2024

Trust Recognition in Human-Robot Cooperation Using EEG

ICRA 2024poster

Collaboration between humans and robots is becoming increasingly crucial in our daily life. In order to accomplish efficient cooperation, trust recognition is vital, empowering robots to predict human behaviors and make trust-aware decisions. Consequently, there is an urgent need for a generalized a…

Cited by 3SourcecodeScholar
2024

Ultrafast capturing in-flight objects with reprogrammable working speed ranges

ICRA 2024poster

In-flight high-speed object capturing is crucial in nature to improve survival and adaptation to the environment, such as the predation of frogs, leopards, and eagles. Despite its ubiquitousness in nature, capturing fast-moving objects is extremely challenging in engineering implementations. In this…

Cited by 0SourceScholar
2024

X-Tacformer : Spatio-tempral Attention Model for Tactile Recognition

ICRA 2024poster

Recently, tactile sensing has attracted great interests in robotics, especially for exploring unstructured objects. Sensor arrays play an important role in the exploration, which generates rich spatio-temporal information. In this work, we propose an efficient tactile recognition model, X-Tacformer.…

Cited by 0SourceScholar
2023

Interpretable Motion Planner for Urban Driving via Hierarchical Imitation Learning

IROS 2023poster

Learning-based approaches have achieved remarkable performance in the domain of autonomous driving. Leveraging the impressive ability of neural networks and large amounts of human driving data, complex patterns and rules of driving behavior can be encoded as a model to benefit the autonomous driving…

Cited by 5SourceScholar
2022

A Novel Neural Multi-Store Memory Network for Autonomous Visual Navigation in Unknown Environment

RA-L 2022

Learning to achieve a user-specified objective from a random position in unseen environments is challenging for image-guided navigation agents. The abilities of long-horizon reasoning and semantic understanding are still lacking. Inspired by the human memory mechanism, we introduce a neural multi-st

Cited by 24SourceScholar
2021

Domain-Smoothing Network for Zero-Shot Sketch-Based Image Retrieval

IJCAI 2021poster

Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a novel cross-modal retrieval task, where abstract sketches are used as queries to retrieve natural images under zero-shot scenario. Most existing methods regard ZS-SBIR as a traditional classification problem and employ a cross-entropy or triplet-…

2018

Scale Aggregation Network for Accurate and Efficient Crowd Counting

ECCV 2018poster

In this paper, we propose a novel encoder-decoder network, called extit{Scale Aggregation Network (SANet)}, for accurate and efficient crowd counting. The encoder extracts multi-scale features with scale aggregation modules and the decoder generates high-resolution density maps by using a set of tra…

Cited by 835SourcePDFScholar