← Search

Zhi Gao

34 accepted papers

2026

FAR-RIO: A Fast and Robust Radar-Inertial Odometry With Isotropic Uncertainty Model and Dual-Observation Update Pipeline

RA-L 2026

Due to the ability to provide point clouds and Doppler velocity, as well as the adaptability in harsh weather conditions, 4D Radar has emerged as a new option for Simultaneous Localization and Mapping (SLAM). However, there is limited research on both robustness and computational efficiency, which a

Cited by 0SourceScholar
2026

KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Controls

ICML 2026poster

Large Multimodal Models encode extensive factual knowledge in their pre-trained weights. However, its knowledge remains static and limited, unable to keep pace with real-world developments, which hinders continuous knowledge acquisition. Effective knowledge injection thus becomes critical, involving…

Cited by 0SourceScholar
2026

Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds

ICLR 2026poster

Modality alignment is critical for vision-language models (VLMs) to effectively integrate information across modalities. However, existing methods extract hierarchical features from text while representing each image with a single feature, leading to asymmetric and suboptimal alignment. To address t…

Cited by 0SourceScholar
2026

STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning

ICLR 2026poster

In vision–language models (VLMs), misalignment between textual descriptions and visual coordinates often induces hallucinations. This issue becomes particularly severe in dense prediction tasks such as spatial–temporal video grounding (STVG). Prior approaches typically focus on enhancing visual–text…

Cited by 0SourceScholar
2026

TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents

AAAI 2026technical

Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. However, a major challenge in developing generalized GUI agents is the lack of sufficient trajectory data across various o

Cited by 0SourcePDFScholar
2026

VUDG: A Dataset for Video Understanding Domain Generalization

ICLR 2026poster

Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets. However, the robustness of these models to domain shifts encountered in real-world video applications remains a critical yet underexplor…

Cited by 0SourceScholar
2026

When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations

ICLR 2026poster

Large Multimodal Models (LMMs) store vast amounts of pretrained knowledge but struggle to remain aligned with real-world updates, making it difficult to avoid capability degradation when acquiring evolving knowledge. Furthermore, most current work focuses on exploring static textual knowledge inject…

Cited by 0SourceScholar
2025

Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning

NeurIPS 2025poster

Open-vocabulary learning requires modeling the data distribution in open environments, which consists of both seen-class and unseen-class data. Existing methods estimate the distribution in open environments using seen-class data, where the absence of unseen classes makes the estimation error inhe…

Cited by 0SourceScholar
2025

Enhancing the Utilization of Color Information in Point Cloud Semantic Segmentation

ICRA 2025

Point cloud semantic segmentation is crucial in various applications such as autonomous driving, robotics, and virtual reality, aiming to assign labels to each point in a cloud to reflect spatial relationships and boundaries. While previous methods primarily focus on geometric features, they often o

Cited by 0SourceScholar
2025

Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning

NeurIPS 2025poster

Multimodal agents, which integrate a controller (e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks. Existing approaches for training these agents, both supervised fine-tuning and reinforcement learning, depend on extens…

Cited by 0SourceScholar
2025

MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge

ICLR 2025poster

Knowledge editing techniques have emerged as essential tools for updating the factual knowledge of large language models (LLMs) and multimodal models (LMMs), allowing them to correct outdated or inaccurate information without retraining from scratch. However, existing benchmarks for multimodal knowl…

2025

Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

ICLR 2025spotlight

The advancement of large language models (LLMs) prompts the development of multi-modal agents, which are used as a controller to call external tools, providing a feasible way to solve practical tasks. In this paper, we propose a multi-modal agent tuning method that automatically generates multi-moda…

Cited by 5SourcePDFScholar
2024

Accurate and Efficient Loop Closure Detection With Deep Binary Image Descriptor and Augmented Point Cloud Registration

IROS 2024poster

Loop Closure Detection (LCD) is an essential component of Simultaneous Localization and Mapping (SLAM), helping to correct drift errors, facilitate map merging, or both by identifying previously observed scenes. Despite its importance, traditional LCD algorithms based on single sensor such as camera…

Cited by 0SourceScholar
2024

CLOVA: A Closed-LOop Visual Assistant with Tool Usage and Update

CVPR 2024poster

Utilizing large language models (LLMs) to compose off-the-shelf visual tools represents a promising avenue of research for developing robust visual assistants capable of addressing diverse visual tasks. However these methods often overlook the potential for continual learning typically by freezing t…

Cited by 29SourcePDFScholar
2024

EnYOLO: A Real-Time Framework for Domain-Adaptive Underwater Object Detection with Image Enhancement

ICRA 2024poster

In recent years, significant progress has been made in the field of underwater image enhancement (UIE). However, its practical utility for high-level vision tasks, such as underwater object detection (UOD) in Autonomous Underwater Vehicles (AUVs), remains relatively unexplored. It may be attributed…

Cited by 2SourceScholar
2024

FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal Models

NeurIPS 2024poster

Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spon…

Cited by 4SourcePDFScholar
2024

SGCalib: A Two-stage Camera-LiDAR Calibration Method Using Semantic Information and Geometric Features

ICRA 2024poster

Extrinsic calibration is an essential prerequisite for the applications of camera-LiDAR fusion. Existing methods either suffer from the complex offline setting of man-made targets or tend to produce suboptimal and unrobust results. In this paper, we propose an online two-stage calibration method tha…

Cited by 4SourceScholar
2024

VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

ECCV 2024poster

"We explore how reconciling several foundation models (large language models and vision-language models) with a novel unified memory mechanism could tackle the challenging video understanding problem, especially capturing the long-term temporal relations in lengthy videos. In particular, the propose…

2023

SyreaNet: A Physically Guided Underwater Image Enhancement Framework Integrating Synthetic and Real Images

ICRA 2023poster

Underwater image enhancement (UIE) is vital for high-level vision-related underwater tasks. Although learning-based UIE methods have made remarkable achievements in recent years, it's still challenging for them to consistently deal with various underwater conditions, which could be caused by: 1) the…

Cited by 37SourcecodeScholar
2023

TJ-FlyingFish: Design and Implementation of an Aerial-Aquatic Quadrotor with Tiltable Propulsion Units

ICRA 2023poster

Aerial-aquatic vehicles are capable to move in the two most dominant fluids, making them more promising for a wide range of applications. We propose a prototype with special designs for propulsion and thruster configuration to cope with the vast differences in the fluid properties of water and air.…

Cited by 34SourceScholar
2022

Efficient Riemannian Meta-Optimization by Implicit Differentiation

AAAI 2022technical

To solve optimization problems with nonlinear constrains, the recently developed Riemannian meta-optimization methods show promise, which train neural networks as an optimizer to perform optimization on Riemannian manifolds. A key challenge is the heavy computational and memory burdens, because com…

2022

Hyperbolic Feature Augmentation via Distribution Estimation and Infinite Sampling on Manifolds

NeurIPS 2022accept

Learning in hyperbolic spaces has attracted growing attention recently, owing to their capabilities in capturing hierarchical structures of data. However, existing learning algorithms in the hyperbolic space tend to overfit when limited data is given. In this paper, we propose a hyperbolic feature a…

Cited by 12SourcePDFScholar
2022

WeakLabel3D-Net: A Complete Framework for Real-Scene LiDAR Point Clouds Weakly Supervised Multi-Tasks Understanding

ICRA 2022poster

Existing state-of-the-art 3D point clouds understanding methods only perform well in a fully supervised manner. To the best of our knowledge, there exists no unified framework which simultaneously solves the downstream high-level understanding tasks, especially when labels are extremely limited. Thi…

Cited by 31SourceScholar
2022

Weakly Supervised 3D Scene Segmentation with Region-Level Boundary Awareness and Instance Discrimination

ECCV 2022poster

"Current state-of-the-art 3D scene understanding methods are merely designed in a full-supervised way. However, in the limited reconstruction cases, only limited 3D scenes can be reconstructed and annotated. We are in need of a framework that can concurrently be applied to 3D point cloud semantic se…

Cited by 48SourcePDFScholar
2021

FG-Conv: Large-Scale LiDAR Point Clouds Understanding Leveraging Feature Correlation Mining and Geometric-Aware Modeling

ICRA 2021poster

This work presents a general deep learning framework for large-scale point clouds understanding without voxelizations, called FG-Conv, which achieves an accurate and real-time understanding of point clouds. Through our novel design combining feature level correlation mining and deformable convolutio…

Cited by 31SourceScholar
2021

Learning a Gradient-free Riemannian Optimizer on Tangent Spaces

AAAI 2021technical

A principal way of addressing constrained optimization problems is to model them as problems on Riemannian manifolds. Recently, Riemannian meta-optimization provides a promising way for solving constrained optimization problems by learning optimizers on Riemannian manifolds in a data-driven fashion,…

2020

A Target Tracking and Positioning Framework for Video Satellites Based on SLAM

IROS 2020poster

With the booming development in aerospace technology, the video satellite which observes the live phenomena on the ground by video shooting has gradually emerged as a new Earth observation method. And remote sensing comes into a "dynamic" era with the demand for new processing techniques, especially…

Cited by 5SourceScholar