← Search

Yunzhi Lin

11 accepted papers

2025

OmniPose6D: Towards Short-Term Object Pose Tracking in Dynamic Scenes from Monocular RGB

IROS 2025

To address the challenge of short-term object pose tracking in dynamic environments with monocular RGB input, we introduce a large-scale synthetic dataset Omni-Pose6D, crafted to mirror the diversity of real-world conditions. We additionally present a benchmarking framework for a comprehensive compa

Cited by 1SourceScholar
2023

KGNv2: Separating Scale and Pose Prediction for Keypoint-Based 6-DoF Grasp Synthesis on RGB-D Input

IROS 2023poster

We propose an improved keypoint approach for 6-DoF grasp pose synthesis from RGB-D input. Keypoint-based grasp detection from image input demonstrated promising results in a previous study, where the visual information provided by color imagery compensates for noisy or imprecise depth measurements.…

Cited by 4SourcecodeScholar
2023

Keypoint-GraspNet: Keypoint-based 6-DoF Grasp Generation from the Monocular RGB-D input

ICRA 2023poster

The success of 6-DoF grasp learning with point cloud input is tempered by the computational costs resulting from their unordered nature and pre-processing needs for reducing the point cloud to a manageable size. These properties lead to failure on small objects with low point cloud cardinality. Inst…

Cited by 13SourcecodeScholar
2023

Parallel Inversion of Neural Radiance Fields for Robust Pose Estimation

ICRA 2023poster

We present a parallelized optimization method based on fast Neural Radiance Fields (NeRF) for estimating 6-DoF pose of a camera with respect to an object or scene. Given a single observed RGB image of the target, we can predict the translation and rotation of the camera by minimizing the residual be…

Cited by 75SourcecodeScholar
2023

WDiscOOD: Out-of-Distribution Detection via Whitened Linear Discriminant Analysis

ICCV 2023poster

Deep neural networks are susceptible to generating overconfident yet erroneous predictions when presented with data beyond known concepts. This challenge underscores the importance of detecting out-of-distribution (OOD) samples in the open world. In this work, we propose a novel feature-space OOD de…

Cited by 7PDFcodeScholar
2022

Keypoint-Based Category-Level Object Pose Tracking from an RGB Sequence with Uncertainty Estimation

ICRA 2022poster

We propose a single-stage, category-level 6-DoF pose estimation algorithm that simultaneously detects and tracks instances of objects within a known category. Our method takes as input the previous and current frame from a monocular RGB video, as well as predictions from the previous frame, to predi…

Cited by 29SourceScholar
2022

SGL: Symbolic Goal Learning in a Hybrid, Modular Framework for Human Instruction Following

RA-L 2022

This paper investigates human instruction following for robotic manipulation via a hybrid, modular system with symbolic and connectionist elements. Symbolic methods build modular systems with semantic parsing and task planning modules for producing sequences of actions from natural language requests

Cited by 7SourcecodeScholar
2022

Single-Stage Keypoint- Based Category-Level Object Pose Estimation from an RGB Image

ICRA 2022poster

Prior work on 6-DoF object pose estimation has largely focused on instance-level processing, in which a textured CAD model is available for each object being detected. Category-level 6- DoF pose estimation represents an important step toward developing robotic vision systems that operate in unstruct…

Cited by 63SourcecodeScholar
2021

A Joint Network for Grasp Detection Conditioned on Natural Language Commands

ICRA 2021poster

We consider the task of grasping a target object based on a natural language command query. Previous work primarily focused on localizing the object given the query, which requires a separate grasp detection module to grasp it. The cascaded application of two pipelines incurs errors in overlapping m…

Cited by 54SourceScholar
2021

Multi-view Fusion for Multi-level Robotic Scene Understanding

IROS 2021poster

We present a system for multi-level scene awareness for robotic manipulation. Given a sequence of camera-inhand RGB images, the system calculates three types of information: 1) a point cloud representation of all the surfaces in the scene, for the purpose of obstacle avoidance. 2) the rough pose of…

Cited by 37SourceScholar
2020

Using Synthetic Data and Deep Networks to Recognize Primitive Shapes for Object Grasping

ICRA 2020poster

A segmentation-based architecture is proposed to decompose objects into multiple primitive shapes from monocular depth input for robotic manipulation. The backbone deep network is trained on synthetic data with 6 classes of primitive shapes generated by a simulation engine. Each primitive shape is d…

Cited by 54SourceScholar