← Search

Irving Fang

6 accepted papers

2025

Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction

ICRA 2025

Humans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense, a novel 3D reconstruction framework that enables robots to fuse priors from foundation models with highly sparse observ

Cited by 6SourceScholar
2025

GARF: Learning Generalizable 3D Reassembly for Real-World Fractures

ICCV 2025poster

3D reassembly is a challenging spatial intelligence task with broad applications across scientific domains. While large-scale synthetic datasets have fueled promising learning-based approaches, their generalizability to different domains is limited. Critically, it remains uncertain whether models tr…

Cited by 0SourcePDFScholar
2025

VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model

IROS 2025

Large Vision Language Models (VLMs) have been adopted in robotics for their strong common sense understanding and generalization capabilities. Existing works leverage VLMs for task and motion planning based on language instructions and robot observations. In this work, we explore using VLM to interp

Cited by 45SourcecodeScholar
2024

EgoPAT3Dv2: Predicting 3D Action Target from 2D Egocentric Vision for Human-Robot Interaction

ICRA 2024poster

A robot’s ability to anticipate the 3D action target location of a hand’s movement from egocentric videos can greatly improve safety and efficiency in human-robot interaction (HRI). While previous research predominantly focused on semantic action classification or 2D target region prediction, we arg…

Cited by 2SourceScholar
2024

LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images

CVPR 2024highlight

Lithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material which is critical for understanding archaeological artifacts material interactions tool functionalities and dental records. However this challeng…

Cited by 4SourcePDFScholar
2023

Metric-Free Exploration for Topological Mapping by Task and Motion Imitation in Feature Space

RSS 2023poster

We propose DeepExplorer, a simple and lightweight metric-free exploration method for topological mapping of unknown environments. It performs task and motion planning (TAMP) entirely in image feature space. The task planner is a recurrent network using the latest image observation sequence to halluc…