← Search

Zhe Huang

20 accepted papers

2026

AerialExtreMatch: A Benchmark for Extreme-View Image Matching and Localization

RA-L 2026

Image matching serves as a core component for UAV localization guided by satellite imagery. However, this task remains highly challenging due to the extreme viewpoint discrepancies between low-altitude UAV images and nadir-view satellite maps. Existing datasets primarily focus on ground-level or hig

Cited by 0SourcecodeScholar
2026

LacTokGen: Latent Consistency Tokenizer for 1024-pixel Image Generation by 256 Tokens

CVPR 2026

Image tokenization has significantly advanced visual generation and multimodal modeling, particularly when paired with autoregressive models. However, current methods face challenges in balancing efficiency and quality: high-resolution image generation either requires an excessive number of tokens o

Cited by 0SourcecodeScholar
2025

MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing

CVPR 2025poster

Deep visual odometry has demonstrated great advancements by learning-to-optimize technology. This approach heavily relies on the visual matching across frames. However, ambiguous matching in challenging scenarios leads to significant errors in geometric modeling and bundle adjustment optimization,…

Cited by 0SourcePDFScholar
2025

Towards Real-Time Generation of Delay-Compensated Video Feeds for Outdoor Mobile Robot Teleoperation

ICRA 2025

Teleoperation is an important technology to enable supervisors to control agricultural robots remotely. However, environmental factors in dense crop rows and limitations in network infrastructure hinder the reliability of data streamed to teleoperators. These issues result in delayed and variable fr

Cited by 3SourceScholar
2024

InterLUDE: Interactions between Labeled and Unlabeled Data to Enhance Semi-Supervised Learning

ICML 2024poster

Semi-supervised learning (SSL) seeks to enhance task performance by training on both labeled and unlabeled data. Mainstream SSL image classification methods mostly optimize a loss that additively combines a supervised classification objective with a regularization term derived *solely* from unlabele…

2024

Multi-Granular Transformer for Motion Prediction with LiDAR

ICRA 2024poster

Motion prediction has been an essential component of autonomous driving systems since it handles highly uncertain and complex scenarios involving moving agents of different types. In this paper, we propose a Multi-Granular TRansformer (MGTR) framework, an encoder-decoder network that exploits contex…

Cited by 11SourceScholar
2024

Neural Informed RRT*: Learning-based Path Planning with Point Cloud State Representations under Admissible Ellipsoidal Constraints

ICRA 2024poster

Sampling-based planning algorithms like Rapidly-exploring Random Tree (RRT) are versatile in solving path planning problems. RRT* offers asymptotic optimality but requires growing the tree uniformly over the free space, which leaves room for efficiency improvement. To accelerate convergence, rule-ba…

Cited by 16SourcecodeScholar
2024

Split-and-Denoise: Protect large language model inference with local differential privacy

ICML 2024poster

Large Language Models (LLMs) excel in natural language understanding by capturing hidden semantics in vector space. This process enriches the value of text embeddings for various downstream tasks, thereby fostering the Embedding-as-a-Service (EaaS) business model. However, the risk of privacy leakag…

2024

Systematic Comparison of Semi-supervised and Self-supervised Learning for Medical Image Classification

CVPR 2024poster

In typical medical image classification problems labeled data is scarce while unlabeled data is more available. Semi-supervised learning and self-supervised learning are two different research directions that can improve accuracy by learning from extra unlabeled data. Recent methods from both direct…

2024

VOLoc: Visual Place Recognition by Querying Compressed Lidar Map

ICRA 2024poster

The availability of city-scale Lidar maps enables the potential of city-scale place recognition using mobile cameras. However, the city-scale Lidar maps generally need to be compressed for storage efficiency, which increases the difficulty of direct visual place recognition in compressed Lidar maps.…

Cited by 7SourcecodeScholar
2023

Fix-A-Step: Semi-supervised Learning From Uncurated Unlabeled Data

AISTATS 2023poster

Semi-supervised learning (SSL) promises improved accuracy compared to training classifiers on small labeled datasets by also training on many unlabeled images. In real applications like medical imaging, unlabeled data will be collected for expediency and thus uncurated: possibly different from the l…

2023

Hierarchical Intention Tracking for Robust Human-Robot Collaboration in Industrial Assembly Tasks

ICRA 2023poster

Collaborative robots require effective human intention estimation to safely and smoothly work with humans in less structured tasks such as industrial assembly, where human intention continuously changes. We propose the concept of intention tracking and introduce a collaborative robot system that con…

Cited by 16SourceScholar
2023

Intention Aware Robot Crowd Navigation with Attention-Based Interaction Graph

ICRA 2023poster

We study the problem of safe and intention-aware robot navigation in dense and interactive crowds. Most previous reinforcement learning (RL) based methods fail to consider different types of interactions among all agents or ignore the intentions of people, which results in performance degradation. I…

Cited by 92SourceScholar
2022

Addressing Optimism Bias in Sequence Modeling for Reinforcement Learning

ICML 2022spotlight

Impressive results in natural language processing (NLP) based on the Transformer neural network architecture have inspired researchers to explore viewing offline reinforcement learning (RL) as a generic sequence modeling problem. Recent works based on this paradigm have achieved state-of-the-art res…

2022

Learning Sparse Interaction Graphs of Partially Detected Pedestrians for Trajectory Prediction

RA-L 2022

Multi-pedestrian trajectory prediction is an indispensable element of autonomous systems that safely interact with crowds in unstructured environments. Many recent efforts in trajectory prediction algorithms have focused on understanding social norms behind pedestrian motions. Yet we observe these w

Cited by 29SourcecodeScholar
2021

Long-Term Pedestrian Trajectory Prediction Using Mutable Intention Filter and Warp LSTM

RA-L 2021

Trajectory prediction is one of the key capabilities for robots to safely navigate and interact with pedestrians. Critical insights from human intention and behavioral patterns need to be integrated to effectively forecast long-term pedestrian behavior. Thus, we propose a framework incorporating a m

Cited by 38SourcecodeScholar
2021

The Tufts fNIRS Mental Workload Dataset & Benchmark for Brain-Computer Interfaces that Generalize

NeurIPS 2021poster

Functional near-infrared spectroscopy (fNIRS) promises a non-intrusive way to measure real-time brain activity and build responsive brain-computer interfaces. A primary barrier to realizing this technology's potential has been that observed fNIRS signals vary significantly across human users. Buildi…

Cited by 31SourceScholar
2020

3D Electromagnetic Reconfiguration Enabled by Soft Continuum Robots

RA-L 2020

The properties of radio frequency electromagnetic systems can be manipulated by changing the 3D geometry of the system. Most reconfiguration schemes specify different conductive pathways using electrical switches in mechanically static systems, or they actuate and reshape a continuous metallic struc

Cited by 19SourceScholar