← Search

Hongyu Zhou

22 accepted papers

2026

Adaptive Legged Locomotion Via Online Learning for Model Predictive Control

ICRA 2026poster

We provide an algorithm for adaptive legged locomotion via online learning and model predictive control. The algorithm is composed of two interacting modules: model predictive control (MPC) and online learning of residual dynamics. The residual dynamics can represent modeling errors and external dis…

2026

Cross-Distill: Multi-Manifold and Viewpoint-Decoupled Distillation for Cross-View Geo-Localization

ICRA 2026poster

Abstract— Cross-View Geo-Localization (CVGL) localizes a query image via retrieval from georeferenced satellite imagery,yet severe viewpoint variation remains a central challenge. Recent advances often rely on heavy backbones or add-on modules that achieve high accuracy but are impractical on resour…

Cited by 0Scholar
2026

NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

ICLR 2026oral

Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ vector quantization (VQ) to obtain discrete tokens with quantization loss. In this paper, we push the autoregressive parad…

Cited by 0SourcecodeScholar
2026

SAFL-Geo: Structure-Aware Feature Learning with Fusion Loss for Infrared-Visible Geo-Localization

ICRA 2026poster

Cross-modal Visual Geo-localization often aims to retrieve a satellite visible-light image of the same geographic lo cation from a large-scale database using an infrared image cap tured by an unmanned aerial vehicle (UAV), thereby achieving precise localization. This capability is crucial for autono…

Cited by 0Scholar
2025

EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View Synthesis

CVPR 2025poster

Novel view synthesis of urban scenes is essential for autonomous driving-related applications. Existing NeRF and 3DGS-based methods show promising results in achieving photorealistic renderings but require slow, per-scene optimization. We introduce EVolSplat, an efficient 3D Gaussian Splatting model…

Cited by 0SourcePDFScholar
2025

JRN-Geo: A Joint Perception Network Based on RGB and Normal Images for Cross-View Geo-Localization

ICRA 2025

Cross-view geo-localization plays a critical role in Unmanned Aerial Vehicle (UAV) localization and navigation. However, significant challenges arise from the drastic viewpoint differences and appearance variations between images. Existing methods predominantly rely on semantic features from RGB ima

Cited by 1SourceScholar
2025

Learning Heterogeneous Tissues with Mixture of Experts for Gigapixel Whole Slide Images

CVPR 2025poster

Analyzing gigapixel Whole Slide Images (WSIs) is challenging due to the complex pathological tissue environment and the absence of target-driven domain knowledge. Previous methods incorporated pathological priors to mitigate this issue but relied on additional inference steps and specialized workflo…

2025

Learning Temporally Consistent Video Depth from Video Diffusion Priors

CVPR 2025poster

This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate…

2025

UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene Simulation

CVPR 2025poster

Photorealistic 3D vehicle models with high controllability are essential for autonomous driving simulation and data augmentation. While handcrafted CAD models provide flexible controllability, free CAD libraries often lack the high-quality materials necessary for photorealistic rendering. Conversely…

Cited by 0SourcePDFScholar
2024

ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

IJCAI 2024poster

Human-AI interactivity is a critical aspect that reflects the usability of Multimodal Large Language Models (MLLMs). However, existing end-to-end MLLMs only allow users to interact with them through language instructions, leading to the limitation of the interactive accuracy and efficiency. In this…

2024

DreamLLM: Synergistic Multimodal Comprehension and Creation

ICLR 2024spotlight

This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation. DreamLLM operates on two fundamental principles. The first focuses on the generative m…

2024

Exploring Recurrent Long-Term Temporal Fusion for Multi-View 3D Perception

RA-L 2024

Long-term temporal fusion is a crucial but often overlooked technique in camera-based Bird's-Eye-View (BEV) 3D perception. Existing methods are mostly in a parallel manner. While parallel fusion can benefit from long-term information, it suffers from increasing computational and memory overheads as

Cited by 96SourceScholar
2024

HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting

CVPR 2024poster

Holistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel view synthesis parsing semantic labels and tracking moving objects. Despite considerable progress existing approaches often…

2024

OpenStreetView-5M: The Many Roads to Global Visual Geolocation

CVPR 2024poster

Determining the location of an image anywhere on Earth is a complex visual task which makes it particularly relevant for evaluating computer vision algorithms. Determining the location of an image anywhere on Earth is a complex visual task which makes it particularly relevant for evaluating computer…

2023

MatrixVT: Efficient Multi-Camera to BEV Transformation for 3D Perception

ICCV 2023poster

This paper proposes an efficient multi-camera to Bird's-Eye-View (BEV) view transformation method for 3D perception, dubbed MatrixVT. Existing view transformers either suffer from poor transformation efficiency or rely on device-specific operators, hindering the broad application of BEV models. In…

Cited by 47PDFcodeScholar
2023

Online Submodular Coordination With Bounded Tracking Regret: Theory, Algorithm, and Applications to Multi-Robot Coordination

RA-L 2023

We enable efficient and effective coordination in unpredictable environments, i.e., in environments whose future evolution is unknown a priori and even adversarial. We are motivated by the future of autonomy that involves multiple robots coordinating in dynamic, unstructured, and adversarial environ

Cited by 15SourceScholar
2023

Safe Non-Stochastic Control of Control-Affine Systems: An Online Convex Optimization Approach

RA-L 2023

We study how to safely control nonlinear control-affine systems that are corrupted with bounded <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">non-stochastic noise</i> , i.e., noise that is unknown a priori and that is <underline xmlns:mml="http://w

Cited by 16SourceScholar
2022

Dense Teacher: Dense Pseudo-Labels for Semi-Supervised Object Detection

ECCV 2022poster

"To date, the most powerful semi-supervised object detectors (SS-OD) are based on pseudo-boxes, which need a sequence of post-processing with fine-tuned hyper-parameters. In this work, we propose replacing the sparse pseudo-boxes with the dense prediction as a united and straightforward form of pseu…

2022

Grouptron: Dynamic Multi-Scale Graph Convolutional Networks for Group-Aware Dense Crowd Trajectory Forecasting

ICRA 2022poster

Accurate, long-term forecasting of pedestrian trajectories in highly dynamic and interactive scenes is a longstanding challenge. Recent advances in using data-driven approaches have achieved significant improvements in terms of prediction accuracy. However, the lack of group-aware analysis has limit…

Cited by 33SourceScholar
2021

EGO-Swarm: A Fully Autonomous and Decentralized Quadrotor Swarm System in Cluttered Environments

ICRA 2021poster

This paper presents a decentralized and asynchronous systematic solution for multi-robot autonomous navigation in unknown obstacle-rich scenes using merely onboard resources. The planning system is formulated under gradient-based local planning framework, where collision avoidance is achieved by for…

Cited by 197SourcecodeScholar