← Search

Wenbo Ding

46 accepted papers

2026

AVR: Active Vision-Driven Precise Robot Manipulation with Viewpoint and Focal Length Optimization

ICRA 2026poster

Robotic manipulation in complex scenes demands precise perception of task-relevant details, yet fixed or suboptimal viewpoints often impair fine-grained perception and induce occlusions, constraining imitation-learned policies. We present AVR (Active Vision-driven Robotics), a bimanual teleoperation…

2026

CEI: A Unified Interface for Cross-Embodiment Visuomotor Policy Learning in 3D Space

RA-L 2026

Robotic foundation models trained on large-scale manipulation datasets have shown promise in learning generalist policies, but they often overfit to specific viewpoints, robot arms, and especially parallel-jaw grippers due to dataset biases. To address this limitation, we propose Cross-Embodiment In

Cited by 0SourcecodeScholar
2026

FlexiCup: Wireless Multimodal Suction Cup With Dual-Zone Vision-Tactile Sensing

RA-L 2026

Conventional suction cups lack sensing capabilities for contact-aware manipulation in unstructured environments. This paper presents FlexiCup, a multimodal suction cup with wireless electronics that integrate dual-zone vision-tactile sensing. The central zone dynamically switches between vision and

Cited by 0SourceScholar
2026

Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation

CVPR 2026

Vision-Language Models (VLMs), leveraging their powerful visual perception and reasoning capabilities, have been widely applied in Unmanned Aerial Vehicle (UAV) tasks.However, the spatial intelligence capabilities of existing VLMs in UAV scenarios remain largely unexplored, raising concerns about th

Cited by 0SourcecodeScholar
2026

Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization

ICLR 2026poster

Critic-free methods like GRPO reduce memory demands by estimating advantages from multiple rollouts but tend to converge slowly, as critical learning signals are diluted by an abundance of uninformative samples and tokens. To tackle this challenge, we propose the **Dynamic Dual-Level Down-Sampling (…

Cited by 0SourceScholar
2026

MoiréTac: A Dual-Mode Visuotactile Sensor for Multidimensional Perception Using Moiré Pattern Amplification

ICRA 2026poster

Visuotactile sensors typically employ sparse marker arrays that limit spatial resolution and lack clear analytical force-to-image relationships. To solve this problem, we present MoiréTac, a dual-mode sensor that generates dense interference patterns via overlapping micro-gratings within a transpare…

2026

Policy Newton Algorithm in Reproducing Kernel Hilbert Space

ICLR 2026poster

Reinforcement learning (RL) policies represented in Reproducing Kernel Hilbert Spaces (RKHS) offer powerful representational capabilities. While second-order optimization methods like Newton's method demonstrate faster convergence than first-order approaches, current RKHS-based policy optimization r…

Cited by 0SourceScholar
2026

SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling

ICLR 2026poster

Training expressive flow-based policies with off-policy reinforcement learning is notoriously unstable due to gradient pathologies in the multi-step action sampling process. We trace this instability to a fundamental connection: the flow rollout is algebraically equivalent to a residual recurrent co…

Cited by 0SourcecodeScholar
2026

What You See Is What You Reach: Towards Spatial Navigation with High-Level Human Instructions

AAAI 2026technical

Embodied navigation is a fundamental capability that enables embodied agents to effectively interact with the physical world in various complex environments. However, a significant gap remains between current embodied navigation tasks and real-world requirements, as existing methods often struggle t

Cited by 0SourcePDFScholar
2025

AirTouch: A Low-Cost Versatile Visuotactile Feedback System for Enhanced Robotic Teleoperation

IROS 2025

Vision-based teleoperation systems are widely used due to their cost-effectiveness and intuitive operation. However, these systems often suffer from challenges such as hand occlusions, environmental variability, and the lack of tactile feedback, limiting their precision and applicability in complex

Cited by 0SourceScholar
2025

Chemistry3D: Robotic Interaction Toolkit for Chemistry Experiments

ICRA 2025

The advent of simulation engines has revolutionized learning and operational efficiency for robots, offering cost-effective and swift pipelines. However, the lack of a universal simulation platform tailored for chemical scenarios impedes progress in robotic manipulation and visualization of reaction

Cited by 2SourcecodeScholar
2025

Depth Restoration of Hand-Held Transparent Objects for Human-to-Robot Handover

ICRA 2025

Transparent objects are common in daily life, while their optical properties pose challenges for RGB-D cameras to capture accurate depth information. This issue is further amplified when these objects are hand-held, as hand occlusions further complicate depth estimation. For assistant robots, howeve

Cited by 3SourceScholar
2025

Distributional Decision Transformer: Risk-Sensitive Offline RL via Quantile-Based Critics and Stochastic Return

IROS 2025

Offline reinforcement learning faces a critical challenge in synthesizing high-reward trajectories from suboptimal datasets while robustly handling the stochasticity inherent in real-world decision-making. While combination of return-conditioned sequence models, such as Decision Transformers (DT), a

Cited by 0SourceScholar
2025

Efficient and Hardware-Friendly Online Adaptation for Deep Stereo Depth Estimation on Embedded Robots

RA-L 2025

Accurate and real-time stereo depth estimation is important for autonomous robots, such as autonomous aerial vehicles (AAVs). Due to the computation constraints of these miniaturized robots, current state-of-the-art algorithms deploy light-weight neural networks while using self-supervised online ad

Cited by 3SourceScholar
2025

Exo-ViHa: A Cross-Platform Exoskeleton System with Visual and Haptic Feedback for Efficient Dexterous Skill Learning

IROS 2025

Imitation learning has emerged as a powerful paradigm for robot skills learning. However, traditional data collection systems for dexterous manipulation face challenges, including a lack of balance between acquisition efficiency, consistency, and accuracy. To address these issues, we introduce Exo-V

Cited by 3SourcecodeScholar
2025

ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model

IROS 2025

Multi-task robotic bimanual manipulation is becoming increasingly popular as it enables sophisticated tasks that require diverse dual-arm collaboration patterns. Compared to unimanual manipulation, bimanual tasks pose challenges to understanding the multi-body spatiotemporal dynamics. An existing me

Cited by 9SourcecodeScholar
2025

Mastering Multi-Drone Volleyball through Hierarchical Co-Self-Play Reinforcement Learning

CoRL 2025poster

In this paper, we tackle the problem of learning to play 3v3 multi-drone volleyball, a new embodied competitive task that requires both high-level strategic coordination and low-level agile control. The task is turn-based, multi-agent, and physically grounded, posing significant challenges due to it…

Cited by 0SourceScholar
2025

Mean-Field Aided QMIX: A Scalable and Flexible Q-Learning Approach for Large-Scale Agent Groups

ICASSP 2025accepted

Value decomposition methods are effective for multi-agent reinforcement learning (MARL), with QMIX being one of the most advanced. However, it struggles with scalability and flexibility in large-scale agent systems. The introduction of mean-field theory into MARL provides a potential solution to bot…

Cited by 0SourceScholar
2025

MonoLDP: LED Assisted Indoor Mobile Bot Monocular Depth Prediction and Pose Estimation System

ICRA 2025

Multi-robot clusters are increasingly deployed in indoor environments, where effective communication and 3D perception are critical for coordinated operations. Monocular cameras, known for their lightweight design, cost-effectiveness, and versatility, present a promising solution for these tasks. Ho

Cited by 1SourcecodeScholar
2025

MuxHand: A Cost-Effective and Compact Dexterous Robotic Hand Using Time-Division Multiplexing Mechanism

IROS 2025

The number of motors directly influences the dexterity, size, and cost of a robotic hand. In this paper, we present MuxHand, a robotic hand that utilizes a time-division multiplexing motor (TDMM) mechanism. This system enables independent control of 9 cables with just 4 motors, significantly reducin

Cited by 1SourceScholar
2025

PUGS: Zero-Shot Physical Understanding with Gaussian Splatting

ICRA 2025

Current robotic systems can understand the categories and poses of objects well. But understanding physical properties like mass, friction, and hardness, in the wild, remains challenging. We propose a new method that reconstructs 3D objects using the Gaussian splatting representation and predicts va

Cited by 11SourcecodeScholar
2025

Residual Kernel Policy Network: Enhancing Stability and Robustness in RKHS-Based Reinforcement Learning

ICLR 2025poster

Achieving optimal performance in reinforcement learning requires robust policies supported by training processes that ensure both sample efficiency and stability. Modeling the policy in reproducing kernel Hilbert space (RKHS) enables efficient exploration of local optimal solutions. However, the sta…

Cited by 0SourcePDFScholar
2025

UltraTac: Integrated Ultrasound-Augmented Visuotactile Sensor for Enhanced Robotic Perception

IROS 2025

Visuotactile sensors provide high-resolution tactile information but are incapable of perceiving the material features of objects. We present UltraTac, an integrated sensor that combines visuotactile imaging with ultrasound sensing through a coaxial optoacoustic architecture. The design shares struc

Cited by 1SourceScholar
2025

Universal Visuo-Tactile Video Understanding for Embodied Interaction

NeurIPS 2025poster

Tactile perception is essential for embodied agents to understand the physical attributes of objects that cannot be determined through visual inspection alone. While existing methods have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate…

Cited by 0SourceScholar
2025

VET: A Visual-Electronic Tactile System for Immersive Human-Machine Interaction

IROS 2025

In the pursuit of deeper immersion in human-machine interaction, achieving higher-dimensional tactile input and output on a single interface has become a key research focus. This study introduces the Visual-Electronic Tactile (VET) System, which builds upon vision-based tactile sensors (VBTS) and in

Cited by 0SourceScholar
2025

VolleyBots: A Testbed for Multi-Drone Volleyball Game Combining Motion Control and Strategic Play

NeurIPS 2025poster

Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we present VolleyBots, a novel robot sports testbed where multiple drones cooperate and compete in the sport of volleybal…

Cited by 0SourcecodeScholar
2025

pFedGPA: Diffusion-based Generative Parameter Aggregation for Personalized Federated Learning

AAAI 2025technical

Federated Learning (FL) offers a decentralized approach to model training, where data remains local and only model parameters are shared between the clients and the central server. Traditional methods, such as Federated Averaging (FedAvg), linearly aggregate these parameters which are usually traine…

Cited by 0SourcePDFScholar
2024

CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video Grounding

AAAI 2024technical

This paper studies the spatio-temporal video grounding task, which aims to localize a spatio-temporal tube in an untrimmed video based on the given text description of an event. Existing one-stage approaches suffer from insufficient space-time interaction in two aspects: i) less precise prediction o…

Cited by 1SourcePDFScholar
2024

DeformNet: Latent Space Modeling and Dynamics Prediction for Deformable Object Manipulation

ICRA 2024poster

Manipulating deformable objects is a ubiquitous task in household environments, demanding adequate representation and accurate dynamics prediction due to the objects’ infinite degrees of freedom. This work proposes DeformNet, which utilizes latent space modeling with a learned 3D representation mode…

Cited by 6SourceScholar
2024

Dual-modal Tactile E-skin: Enabling Bidirectional Human-Robot Interaction via Integrated Tactile Perception and Feedback

ICRA 2024poster

To foster an immersive and natural human-robot interaction (HRI), the implementation of tactile perception and feedback becomes imperative, effectively bridging the conventional sensory gap. In this paper, we propose a dual-modal electronic skin (e-skin) that integrates magnetic tactile sensing and…

Cited by 2SourceScholar
2024

M3ARL: Moment-Embedded Mean-Field Multi-Agent Reinforcement Learning for Continuous Action Space

ICASSP 2024accepted

Mean-field theory offers a promising solution to the scalability issues encountered in multi-agent reinforcement learning (MARL) within large-scale systems. However, most existing MARL algorithms based on mean-field theory are typically constrained to discrete action space. In continuous action spac…

Cited by 0SourceScholar
2024

Optimizing Trading Strategies in Quantitative Markets Using Multi-Agent Reinforcement Learning

ICASSP 2024accepted

Quantitative markets are characterized by swift dynamics and abundant uncertainties, making the pursuit of profit-driven stock trading actions inherently challenging. Within this context, Reinforcement Learning (RL) — which operates on a reward-centric mechanism for optimal control — has surfaced as…

Cited by 0SourceScholar
2024

Point-Wise Vibration Pattern Production via a Sparse Actuator Array for Surface Tactile Feedback

ICRA 2024poster

Surface vibration tactile feedback is capable of conveying various semantic information to humans via handheld electronic devices, such as smartphones, touch panels, and game controllers. However, covering the entire contacting surface of the device with a dense arrangement of actuators can affect i…

Cited by 0SourcecodeScholar
2024

SATac: A Thermoluminescence Enabled Tactile Sensor for Concurrent Perception of Temperature, Pressure, and Shear

ICRA 2024poster

Most vision-based tactile sensors use elastomer deformation to infer tactile information, which can not sense some modalities, like temperature. As an important part of human tactile perception, temperature sensing can help robots better interact with the environment. In this work, we propose a nove…

Cited by 1SourceScholar
2024

VibroBot: A Lightweight and Wirelessly Programmable Vibration Bot for Haptic Guidance

RA-L 2024

Cutaneous haptics is helpful to tame the human-machine mismatch by interactive tactile feedback and perform precise manipulations for virtual immersive interactions. However, wearable tactile gloves cover the palm and fingers greatly, thus reducing the tactile information from interactive objects an

Cited by 3SourceScholar
2023

Autonomous Swarm Robot Coordination via Mean-Field Control Embedding Multi-Agent Reinforcement Learning

IROS 2023poster

The learning approaches of designing a controller to guide the collective behavior of swarm robots have gained significant attention in recent years. However, the scalability of swarm robots and their inherent stochasticity complicate the control problem due to increasing complexity, unpredictabilit…

Cited by 4SourceScholar
2023

Every Parameter Matters: Ensuring the Convergence of Federated Learning with Dynamic Heterogeneous Models Reduction

NeurIPS 2023poster

Cross-device Federated Learning (FL) faces significant challenges where low-end clients that could potentially make unique contributions are excluded from training large models due to their resource bottlenecks. Recent research efforts have focused on model-heterogeneous FL, by extracting reduced-si…

Cited by 38SourcePDFScholar
2023

STEV: Stretchable Triboelectric E-skin enabled Proprioceptive Vibration Sensing for Soft Robot

ICRA 2023poster

Vibration perception is essential for robotic sensing and dynamic control. Nevertheless, due to the rigorous demand for sensor conformability and stretchability, enabling soft robots with proprioceptive vibration sensing remains challenging. This paper proposes a novel liquid metal-based stretchable…

Cited by 6SourceScholar
2023

TGF-Net: Sim2Real Transparent Object 6D Pose Estimation Based on Geometric Fusion

RA-L 2023

Transparent objects are a common part of daily life, but their unique optical properties make estimating their 6D pose a challenging task. In this letter, we propose TGF-Net, a monocular instance-level 6D pose estimation method for transparent objects based on geometric fusion. TGF-Net learns the ed

Cited by 20SourceScholar
2023

Understanding the Robustness of 3D Object Detection With Bird's-Eye-View Representations in Autonomous Driving

CVPR 2023poster

3D object detection is an essential perception task in autonomous driving to understand the environments. The Bird's-Eye-View (BEV) representations have significantly improved the performance of 3D detectors with camera inputs on popular benchmarks. However, there still lacks a systematic understand…

2023

Visuotactile Sensor Enabled Pneumatic Device Towards Compliant Oropharyngeal Swab Sampling

IROS 2023poster

Manual oropharyngeal (OP) swab sampling is an intensive and risky task. In this article, a novel OP swab sampling device of low cost and high compliance is designed by combining the visuotactile sensor and the pneumatic actuator-based gripper. Here, a concave visuotactile sensor called CoTac is firs…

Cited by 3SourceScholar
2023

WS-3D-Lane: Weakly Supervised 3D Lane Detection With 2D Lane Labels

ICRA 2023poster

Compared to 2D lanes, real 3D lane data is difficult to collect accurately. In this paper, we propose a novel method for training 3D lanes with only 2D lane labels, called weakly supervised 3D lane detection WS-3D-Lane. By assumptions of constant lane width and equal height on adjacent lanes, we ind…

Cited by 15SourcecodeScholar
2022

Planar Magnetic Actuation for Soft and Rigid Robots Using a Scalable Electromagnet Array

RA-L 2022

Magnetic actuation system manipulates micro soft or rigid robots by a controllable magnetic field to move them freely in the narrow or enclosed space, which has demonstrated its huge potential in medical interventional surgery and drug delivery. However, the limited working space of paired or area-c

Cited by 16SourceScholar