← Search

Jiuguang Wang

15 accepted papers

2026

CuriousBot: Interactive Mobile Exploration via Actionable 3D Relational Object Graph

RA-L 2026

Mobile exploration is a longstanding challenge in robotics, yet current methods primarily focus on active perception instead of active interaction, limiting the robot's ability to interact with and fully explore its environment. Existing robotic exploration approaches via active interaction are ofte

Cited by 6SourcecodeScholar
2026

From Pixels to Predicates: Learning Symbolic World Models via Pretrained VLMs

RA-L 2026

Our aim is to learn to solve long-horizon decision-making problems in complex robotics domains given low-level skills and a handful of demonstrations containing sequences of images. To this end, we focus on learning abstract symbolic world models that facilitate zero-shot generalization to novel goa

Cited by 0SourceScholar
2026

Generative Models From and for Sampling-Based MPC: A Bootstrapped Approach for Adaptive Contact-Rich Manipulation

RA-L 2026

We present a generative predictive control (GPC) framework that amortizes sampling-based Model Predictive Control (SPC) by bootstrapping it with conditional flow-matching models trained on SPC control sequences collected in simulation. Unlike prior work relying on iterative refinement or gradient-ba

Cited by 0SourceScholar
2026

Judo: A User-Friendly Open-Source Package for Sampling-Based Model Predictive Control

ICRA 2026poster

Sampling-based model predictive control (MPC) is experiencing a resurgence in robotics following both recent hardware successes and advancements in parallelized physics simulation. However, to build on this progress, the robotics community needs to develop shared tools for prototyping, benchmarking,…

2026

Planning-Guided Diffusion Policy Learning for Contact-Rich Bimanual Object Reorientation

ICRA 2026poster

Contact-rich bimanual manipulation involves precise coordination of two arms to change object states through strategically selected contacts and motions. Due to the inherent complexity of these tasks, acquiring sufficient demonstration data and training policies that generalize to unseen scenarios r…

Cited by 0Scholar
2025

ASHiTA: Automatic Scene-grounded HIerarchical Task Analysis

CVPR 2025poster

While recent work in scene reconstruction and understanding has made strides in grounding natural language to physical 3D environments, it is still challenging to ground abstract, high-level instructions to a 3D scene. High-level instructions might not explicitly invoke semantic elements in the scen…

Cited by 0SourcePDFScholar
2025

Is Linear Feedback on Smoothed Dynamics Sufficient for Stabilizing Contact-Rich Plans?

ICRA 2025

Designing planners and controllers for contact-rich manipulation is extremely challenging as contact violates the smoothness conditions that many gradient-based controller synthesis tools assume. Contact smoothing approximates a non-smooth system with a smooth one, allowing one to use these synthesi

Cited by 10SourceScholar
2025

Physics-Driven Data Generation for Contact-Rich Manipulation via Trajectory Optimization

RSS 2025poster

We present a low-cost data generation pipeline that integrates physics-based simulation, human demonstrations, and model-based planning to efficiently generate large-scale, high-quality datasets for contact-rich robotic manipulation tasks. Starting with a small number of embodiment-flexible human de…

Cited by 3PDFScholar
2025

Should We Learn Contact-Rich Manipulation Policies From Sampling-Based Planners?

RA-L 2025

The tremendous success of behavior cloning (BC) in robotic manipulation has been largely confined to tasks where demonstrations can be effectively collected through human teleoperation. However, demonstrations for contact-rich manipulation tasks that require complex coordination of multiple contacts

Cited by 13SourceScholar
2025

Versatile Loco-Manipulation through Flexible Interlimb Coordination

CoRL 2025oral

The ability to flexibly leverage limbs for loco-manipulation is essential for enabling autonomous robots to operate in unstructured environments. Yet, prior work on loco-manipulation is often constrained to specific tasks or predetermined limb configurations. In this work, we present einforcement Le…

Cited by 0SourceScholar
2024

Continuously Improving Mobile Manipulation with Autonomous Real-World RL

CoRL 2024poster

We present a fully autonomous real-world RL framework for mobile manipulation that can learn policies without extensive instrumentation or human supervision. This is enabled by 1) task-relevant autonomy, which guides exploration towards object interactions and prevents stagnation near goal states, 2…

Cited by 3SourcecodeScholar
2024

Equivariant Diffusion Policy

CoRL 2024poster

Recent work has shown diffusion models are an effective approach to learning the multimodal distributions arising from demonstration data in behavior cloning. However, a drawback of this approach is the need to learn a denoising function, which is significantly more complex than learning an explicit…

Cited by 26SourcecodeScholar
2024

GenDP: 3D Semantic Fields for Category-Level Generalizable Diffusion Policy

CoRL 2024poster

Diffusion-based policies have shown remarkable capability in executing complex robotic manipulation tasks but lack explicit characterization of geometry and semantics, which often limits their ability to generalize to unseen objects and layouts. To enhance the generalization capabilities of Diffusio…

Cited by 14SourcecodeScholar
2024

Jacta: A Versatile Planner for Learning Dexterous and Whole-body Manipulation

CoRL 2024poster

Robotic manipulation is challenging due to discontinuous dynamics, as well as high-dimensional state and action spaces. Data-driven approaches that succeed in manipulation tasks require large amounts of data and expert demonstrations, typically from humans. Existing planners are restricted to specif…

Cited by 11SourcecodeScholar
2024

VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation

ICRA 2024poster

Understanding how humans leverage semantic knowledge to navigate unfamiliar environments and decide where to explore next is pivotal for developing robots capable of human-like search behaviors. We introduce a zero-shot navigation approach, Vision-Language Frontier Maps (VLFM), which is inspired by…

Cited by 97SourcecodeScholar