← Search

RUIYU WANG

9 accepted papers

2026

PALM: Enhanced Generalizability for Local Visuomotor Policies via Perception Alignment

RA-L 2026

Generalizing beyond the training domain in image-based behavior cloning remains challenging. Existing methods address individual axes of generalization, workspace shifts, viewpoint changes, and cross-embodiment transfer, yet they are typically developed in isolation and often rely on complex pipelin

Cited by 1SourceScholar
2025

AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play

NeurIPS 2025spotlight

Search-augmented LLMs often struggle with complex reasoning tasks due to ineffective multi-hop retrieval and limited reasoning ability. We propose AceSearcher, a cooperative self-play framework that trains a single large language model (LLM) to alternate between two roles: a decomposer that breaks d…

Cited by 0SourceScholar
2025

CADMorph: Geometry‑Driven Parametric CAD Editing via a Plan–Generate–Verify Loop

NeurIPS 2025poster

A Computer-Aided Design (CAD) model encodes an object in two coupled forms: a \emph{parametric construction sequence} and its resulting \emph{visible geometric shape}. During iterative design, adjustments to the geometric shape inevitably require synchronized edits to the underlying parametric seque…

Cited by 0SourceScholar
2025

Feature Extractor or Decision Maker: Rethinking the Role of Visual Encoders in Visuomotor Policies

ICRA 2025

An end-to-end (E2E) visuomotor policy is typically treated as a unified whole, but recent approaches using out-of-domain (OOD) data to pretrain the visual encoder have cleanly separated the visual encoder from the network, with the remainder referred to as the policy. We propose Visual Alignment Tes

Cited by 1SourceScholar
2025

MirrorDuo: Reflection-Consistent Visuomotor Learning from Mirrored Demonstration Pairs

CoRL 2025poster

Image-based behaviour cloning leverages demonstrations captured from ubiquitous RGB cameras, enabling impressive visuomotor performance. However, it remains constrained by the cost of collecting sufficiently diverse demonstrations, especially for generalizing across workspace variations. We propose…

Cited by 0SourceScholar
2025

Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models

ICML 2025poster

Creating Computer-Aided Design (CAD) models requires significant expertise and effort. Text-to-CAD, which converts textual descriptions into CAD parametric sequences, is crucial in streamlining this process. Recent studies have utilized ground-truth parametric sequences, known as sequential signals…

Cited by 1SourcePDFScholar
2024

Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation

CoRL 2024poster

In vision-based behaviour cloning (BC), traditional image-level augmentation methods such as pixel shifting enhance in-domain performance but often struggle with visual domain shifts, including distractors, occlusion, and changes in lighting and backgrounds. Conversely, superimposition-based augment…

Cited by 2SourceScholar
2024

How Physics and Background Attributes Impact Video Transformers in Robotic Manipulation: A Case Study on Planar Pushing

IROS 2024poster

As model and dataset sizes continue to scale in robot learning, the need to understand how the composition and properties of a dataset affect model performance becomes increasingly urgent to ensure cost-effective data collection and model performance. In this work, we empirically investigate how phy…

Cited by 1SourceScholar