← Search

Harry Zhang

13 accepted papers

2026

FUSE: Quantifying Uncertainty in Multimodal LLMs by Bayesian Fusing Epistemic and Aleatoric Uncertainty

ICML 2026poster

Multimodal large language models (MLLMs) are playing an increasingly important role across multiple domains. In many applications, such as robotics, it is crucial to quantify the uncertainty in the output of these models. } We develop Fused Uncertainty with Semantic Evidence (FUSE), a probabilistic …

Cited by 0SourceScholar
2026

H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows

ICLR 2026poster

Understanding how humans interact with the surrounding environment, and specifically reasoning about object interactions and affordances, is a critical challenge in computer vision, robotics, and AI. Current approaches often depend on labor-intensive, hand-labeled datasets capturing real-world or si…

Cited by 0SourceScholar
2025

CRISP: Object Pose and Shape Estimation with Test-Time Adaptation

CVPR 2025highlight

We consider the problem of estimating object pose and shape from an RGB-D image. Our first contribution is to introduce CRISP, a category-agnostic object pose and shape estimation pipeline. The pipeline implements an encoder-decoder model for shape estimation. It uses FiLM-conditioning for implicit…

2025

Enhancing Autonomous Navigation by Imaging Hidden Objects Using Single-Photon LiDAR

ICRA 2025

Robust autonomous navigation in environments with limited visibility remains a critical challenge in robotics. We present a novel approach that leverages Non-Line-of-Sight (NLOS) sensing using single-photon LiDAR to improve visibility and enhance autonomous navigation. Our method enables mobile robo

Cited by 12SourcecodeScholar
2025

Max Entropy Moment Kalman Filter for Polynomial Systems with Arbitrary Noise

NeurIPS 2025poster

Designing optimal Bayes filters for nonlinear non-Gaussian systems is a challenging task. The main difficulties are: 1) representing complex beliefs, 2) handling non-Gaussian noise, and 3) marginalizing past states. To address these challenges, we focus on polynomial systems and propose the Max Entr…

Cited by 0SourceScholar
2024

Multi-Model 3D Registration: Finding Multiple Moving Objects in Cluttered Point Clouds

ICRA 2024poster

We investigate a variation of the 3D registration problem, named multi-model 3D registration. In the multi-model registration problem, we are given two point clouds picturing a set of objects at different poses (and possibly including points belonging to the background) and we want to simultaneously…

Cited by 13SourceScholar
2023

FlowBot++: Learning Generalized Articulated Objects Manipulation via Articulation Projection

CoRL 2023poster

Understanding and manipulating articulated objects, such as doors and drawers, is crucial for robots operating in human environments. We wish to develop a system that can learn to articulate novel objects with no prior interaction, after training on other articulated objects. Previous approaches fo…

Cited by 35SourceScholar
2022

Enabling On-Device Training of Speech Recognition Models With Federated Dropout

ICASSP 2022accepted

Federated learning can be used to train machine learning models on the edge on local data that never leave devices, providing privacy by default. This presents a challenge pertaining to the communication and computation costs associated with clients’ devices. These costs are strongly correlated with…

Cited by 0SourceScholar
2022

TAX-Pose: Task-Specific Cross-Pose Estimation for Robot Manipulation

CoRL 2022poster

How do we imbue robots with the ability to efficiently manipulate unseen objects and transfer relevant skills based on demonstrations? End-to-end learning methods often fail to generalize to novel objects or unseen configurations. Instead, we focus on the task-specific pose relationship between rele…

Cited by 61SourceScholar
2021

Robots of the Lost Arc: Self-Supervised Learning to Dynamically Manipulate Fixed-Endpoint Cables

ICRA 2021poster

We explore how high-speed robot arm motions can dynamically manipulate ropes and cables to vault over obstacles, knock objects from pedestals, and weave between obstacles. In this paper, we propose a self-supervised learning framework that enables a UR5 robot to perform these three tasks. The framew…

Cited by 72SourceScholar
2020

Dex-Net AR: Distributed Deep Grasp Planning Using a Commodity Cellphone and Augmented Reality App

ICRA 2020poster

Consumer demand for augmented reality (AR) in mobile phone applications, such as the Apple ARKit. Such applications have potential to expand access to robot grasp planning systems such as Dex-Net. AR apps use structure from motion methods to compute a point cloud from a sequence of RGB images taken…

Cited by 19SourceScholar