← Search

Haozhi Qi

18 accepted papers

2026

Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-The-Wild Human Demonstrations

ICRA 2026poster

Learning multi-fingered robot policies from humans performing daily tasks in natural environments has long been a grand goal in the robotics community. Achieving this would mark significant progress toward generalizable robot manipulation in human environments, as it would reduce the reliance on lab…

2026

Learning Dexterous Manipulation Skills from Imperfect Simulations

ICRA 2026poster

Reinforcement learning and sim-to-real transfer have made significant progress in dexterous manipulation. However, progress remains limited by the difficulty of simulating complex contact dynamics and multisensory signals, especially tactile feedback. In this work, we propose DexScrew, a sim-to-real…

2025

From Simple to Complex Skills: The Case of In-Hand Object Reorientation

ICRA 2025

Learning policies in simulation and transferring them to the real world has become a promising approach in dexterous manipulation. However, bridging the sim-to-real gap for each new task requires substantial human effort, such as careful reward engineering, hyperparameter tuning, and system identifi

Cited by 16SourceScholar
2025

Hand-Object Interaction Pretraining from Videos

ICRA 2025

We present an approach to learn general robot manipulation priors from 3D hand-object interaction trajectories. We build a framework to use in-the-wild videos to generate sensorimotor robot trajectories. We do so by lifting both the human hand and the manipulated object in a shared 3D space and reta

Cited by 46SourcecodeScholar
2025

Learning In-Hand Translation Using Tactile Skin with Shear and Normal Force Sensing

ICRA 2025

Recent progress in reinforcement learning (RL) and tactile sensing has significantly advanced dexterous manipulation. However, these methods often utilize simplified tactile signals due to the gap between tactile simulation and the real world. We introduce a sensor model for tactile skin that enable

Cited by 23SourceScholar
2025

Learning Visuotactile Skills With Two Multifingered Hands

ICRA 2025

Aiming to replicate human-like dexterity, perceptual experiences, and motion patterns, we explore learning from human demonstrations using a bimanual system with multifingered hands and visuotactile data. Two significant challenges exist: the lack of an affordable and accessible teleoperation system

Cited by 119SourcecodeScholar
2024

Lessons from Learning to Spin “Pens”

CoRL 2024poster

In-hand manipulation of pen-like objects is a most basic and important skill in our daily lives, as many tools such as hammers and screwdrivers are similarly shaped. However, current learning-based methods struggle with this task due to a lack of high-quality demonstrations and the significant gap b…

Cited by 16SourcecodeScholar
2023

General In-hand Object Rotation with Vision and Touch

CoRL 2023poster

We introduce Rotateit, a system that enables fingertip-based object rotation along multiple axes by leveraging multimodal sensory inputs. Our system is trained in simulation, where it has access to ground-truth object shapes and physical properties. Then we distill it to operate on realistic yet noi…

Cited by 106SourceScholar
2022

Coupling Vision and Proprioception for Navigation of Legged Robots

CVPR 2022poster

We exploit the complementary strengths of vision and proprioception to develop a point-goal navigation system for legged robots, called VP-Nav. Legged systems are capable of traversing more complex terrain than wheeled robots, but to fully utilize this capability, we need a high-level path planner i…

Cited by 73PDFcodeScholar
2022

In-Hand Object Rotation via Rapid Motor Adaptation

CoRL 2022poster

Generalized in-hand manipulation has long been an unsolved challenge of robotics. As a small step towards this grand goal, we demonstrate how to design and learn a simple adaptive controller to achieve in-hand object rotation using only fingertips. The controller is trained entirely in simulation on…

Cited by 115SourcecodeScholar
2021

Learning Long-term Visual Dynamics with Region Proposal Interaction Networks

ICLR 2021poster

Learning long-term dynamics models is the key to understanding physical common sense. Most existing approaches on learning dynamics from visual input sidestep long-term predictions by resorting to rapid re-planning with short-term models. This not only requires such models to be super accurate but a…

2020

Deep Isometric Learning for Visual Recognition

ICML 2020poster

Initialization, normalization, and skip connections are believed to be three indispensable techniques for training very deep convolutional neural networks and obtaining state-of-the-art performance. This paper shows that deep vanilla ConvNets without normalization nor skip connections can also be tr…

2019

Learning to Reconstruct 3D Manhattan Wireframes From a Single Image

ICCV 2019oral

From a single view of an urban environment, we propose a method to effectively exploit the global structural regularities for obtaining a compact, accurate, and intuitive 3D wireframe representation. Our method trains a single convolutional neural network to simultaneously detect salient junctions a…

Cited by 83PDFcodeScholar
2019

NeurVPS: Neural Vanishing Point Scanning via Conic Convolution

NeurIPS 2019poster

We present a simple yet effective end-to-end trainable deep network with geometry-inspired convolutional operators for detecting vanishing points in images. Traditional convolutional neural networks rely on aggregating edge features and do not have mechanisms to directly exploit the geometric proper…

2017

Fully Convolutional Instance-Aware Semantic Segmentation

CVPR 2017spotlight

We present the first fully convolutional end-to-end solution for instance-aware semantic segmentation task. It inherits all the merits of FCNs for semantic segmentation and instance mask proposal. It performs instance mask prediction and classification jointly. The underlying convolutional represent…

Cited by 1418PDFcodeScholar