← Search

Wentao Yuan

16 accepted papers

2026

GraspGen: A Diffusion-Based Framework for 6-DOF Grasping with On-Generator Training

ICRA 2026poster

Grasping is a fundamental robot skill, yet despite significant research advancements, learning-based 6-DOF grasping approaches are still not turnkey and struggle to generalize across different embodiments and in-the-wild settings. We build upon the recent success on modeling the object-centric grasp…

2026

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

RSS 2026poster

Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision–Language Models (VLMs) offer a promising foundation for general-purpose and interpretable robotic reasoning. Aligning VL…

Cited by 0SourceScholar
2025

AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation

ICLR 2025poster

Robotic manipulation in open-world settings requires not only task execution but also the ability to detect and learn from failures. While recent advances in vision-language models (VLMs) and large language models (LLMs) have improved robots' spatial reasoning and problem-solving abilities, they sti…

2024

Avoid Everything: Model-Free Collision Avoidance with Expert-Guided Fine-Tuning

CoRL 2024poster

The world is full of clutter. In order to operate effectively in uncontrolled, real world spaces, robots must navigate safely by executing tasks around obstacles while in proximity to hazards. Creating safe movement for robotic manipulators remains a long-standing challenge in robotics, particularly…

Cited by 3SourceScholar
2024

Evaluating Robustness of Visual Representations for Object Assembly Task Requiring Spatio-Geometrical Reasoning

ICRA 2024poster

This paper primarily focuses on evaluating and benchmarking the robustness of visual representations in the context of object assembly tasks. Specifically, it investigates the alignment and insertion of objects with geometrical extrusions, commonly referred to as a peg-in-hole task. The accuracy req…

Cited by 2SourceScholar
2024

Manipulate-Anything: Automating Real-World Robots using Vision-Language Models

CoRL 2024poster

Large-scale endeavors like RT-1 and widespread community efforts such as Open-X-Embodiment have contributed to growing the scale of robot demonstration data. However, there is still an opportunity to improve the quality, quantity, and diversity of robot demonstration data. Although vision-language m…

Cited by 39SourcecodeScholar
2024

RoboPoint: A Vision-Language Model for Spatial Affordance Prediction in Robotics

CoRL 2024poster

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot behavior, VLMs struggle to precisely articulate robot actions usin…

Cited by 53SourcecodeScholar
2023

M2T2: Multi-Task Masked Transformer for Object-centric Pick and Place

CoRL 2023poster

With the advent of large language models and large-scale robotic datasets, there has been tremendous progress in high-level decision-making for object manipulation. These generic models are able to interpret complex tasks using language commands, but they often have difficulties generalizing to out-…

Cited by 23SourcecodeScholar
2023

TerrainNet: Visual Modeling of Complex Terrain for High-speed, Off-road Navigation

RSS 2023poster

Effective use of camera-based vision systems is essential for robust performance in autonomous off-road driving, particularly in the high-speed regime. Despite success in structured, on-road settings, current end-to-end approaches for scene prediction have yet to be successfully adapted for complex…

Cited by 62SourcePDFScholar
2022

KD-MVS: Knowledge Distillation Based Self-Supervised Learning for Multi-View Stereo

ECCV 2022poster

"Supervised multi-view stereo (MVS) methods have achieved remarkable progress in terms of reconstruction quality, but suffer from the challenge of collecting large-scale ground-truth depth. In this paper, we propose a novel self-supervised training pipeline for MVS based on knowledge distillation, t…

2022

Sobolev Training for Implicit Neural Representations with Approximated Image Derivatives

ECCV 2022poster

"Recently, Implicit Neural Representations (INRs) parameterized by neural networks have emerged as a powerful and promising tool to represent different kinds of signals due to its continuous, differentiable properties, showing superiorities to classical discretized representations. However, the trai…

2021

SORNet: Spatial Object-Centric Representations for Sequential Manipulation

CoRL 2021oral

Sequential manipulation tasks require a robot to perceive the state of an environment and plan a sequence of actions leading to a desired goal state, where the ability to reason about spatial relationships among object entities from raw sensor inputs is crucial. Prior works relying on explicit state…

Cited by 86SourcecodeScholar
2021

STaR: Self-Supervised Tracking and Reconstruction of Rigid Objects in Motion With Neural Rendering

CVPR 2021poster

We present STaR, a novel method that performs Self-supervised Tracking and Reconstruction of dynamic scenes with rigid motion from multi-view RGB videos without any manual annotation. Recent work has shown that neural networks are surprisingly effective at the task of compressing many views of a sce…

Cited by 172PDFScholar
2021

Self-Supervised Learning on 3D Point Clouds by Learning Discrete Generative Models

CVPR 2021poster

While recent pre-training tasks on 2D images have proven very successful for transfer learning, pre-training for 3D data remains challenging. In this work, we introduce a general method for 3D self-supervised representation learning that 1) remains agnostic to the underlying neural network architect…

Cited by 74PDFScholar
2020

DeepGMR: Learning Latent Gaussian Mixture Models for Registration

ECCV 2020poster

Point cloud registration is a fundamental problem in 3D computer vision, graphics and robotics. For the last few decades, existing registration algorithms have struggled in situations with large transformations, noise, and time constraints. In this paper, we introduce Deep Gaussian Mixture Registrat…

2018

Intelligent Shipwreck Search Using Autonomous Underwater Vehicles

ICRA 2018poster

This paper presents an autonomous robot system that is designed to autonomously search for and geo-localize potential underwater archaeological sites. The system, based on Autonomous Underwater Vehicles, invokes a multi-step pipeline. First, the AUV constructs a high altitude scan over a large area…

Cited by 30SourceScholar