← Search

Yi Ru Wang

9 accepted papers

2026

MolmoAct: Action Reasoning Models That Can Reason in Space

ICRA 2026poster

Reasoning is essential for purposeful action, yet most robotic foundation models map perception and instructions directly to control, limiting adaptability, generalization, and semantic grounding. We introduce Action Reasoning Models (ARMs), which integrate perception, planning, and control through …

2026

RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

ICRA 2026poster

We introduce RoboEval, a structured evaluation framework and benchmark for robotic manipulation that augments binary success with principled behavioral and outcome metrics. Existing evaluations often collapse performance into outcome counts, masking differences in execution quality and obscuring fai…

2025

AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation

ICLR 2025poster

Robotic manipulation in open-world settings requires not only task execution but also the ability to detect and learn from failures. While recent advances in vision-language models (VLMs) and large language models (LLMs) have improved robots' spatial reasoning and problem-solving abilities, they sti…

2025

SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation

ICML 2025poster

Robotic manipulation systems operating in diverse, dynamic environments must exhibit three critical abilities: multitask interaction, generalization to unseen scenarios, and spatial memory. While significant progress has been made in robotic manipulation, existing approaches often fall short in gene…

2024

Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

CoRL 2024poster

Large-scale endeavors like RT-1 and widespread community efforts such as Open-X-Embodiment have contributed to growing the scale of robot demonstration data. However, there is still an opportunity to improve the quality, quantity, and diversity of robot demonstration data. Although vision-language m…

Cited by 39SourcecodeScholar
2023

MVTrans: Multi-View Perception of Transparent Objects

ICRA 2023poster

Transparent object perception is a crucial skill for applications such as robot manipulation in household and laboratory settings. Existing methods utilize RGB-D or stereo inputs to handle a subset of perception tasks including depth and pose estimation. However transparent object perception remains…

Cited by 26SourcecodeScholar
2023

NEWTON: Are Large Language Models Capable of Physical Reasoning?

EMNLP 2023long findings

Large Language Models (LLMs), through their contextualized representations, have been empirically proven to encapsulate syntactic, semantic, word sense, and common-sense knowledge. However, there has been limited exploration of their physical reasoning abilities, specifically concerning the crucial…

Cited by 0SourcecodeScholar
2021

Seeing Glass: Joint Point-Cloud and Depth Completion for Transparent Objects

CoRL 2021oral

The basis of many object manipulation algorithms is RGB-D input. Yet, commodity RGB-D sensors can only provide distorted depth maps for a wide range of transparent objects due light refraction and absorption. To tackle the perception challenges posed by transparent objects, we propose TranspareNet,…

Cited by 62SourceScholar