← Search

Zhiheng Li

42 accepted papers

2026

4DRaL: Bridging 4D Radar with LiDAR for Place Recognition Using Knowledge Distillation

ICRA 2026poster

Place recognition is crucial for loop closure detection and global localization in robotics. Although mainstream algorithms typically rely on cameras and LiDAR, these sensors are susceptible to adverse weather conditions. Fortunately, the recently developed 4D millimeter-wave radar (4D radar) offers…

2026

Diffusion-Enhanced Tree Planning for Autonomous Driving

RA-L 2026

In highly interactive urban driving, decision making is often naturally multi-stage, and decisions at different stages can lead to different reactions from surrounding vehicles. This calls for stage-wise evaluation and selection. Tree-based planning naturally supports multi-stage search and evaluati

Cited by 0SourceScholar
2026

KiGRAS: Kinematic-Driven Generative Model for Realistic Agent Simulation

ICRA 2026poster

Trajectory generation is a pivotal task in autonomous driving. Recent studies have introduced the autoregressive paradigm, leveraging the state transition model to approximate future trajectory distributions. This paradigm closely mirrors the real-world trajectory generation process and has achieved…

2026

Multimodal Meta-Verifier with Explicit Structured Recalibration

ICML 2026poster

Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling generalist foundation models. In this work, we investigate ***multimodal meta-verification***, which leverages verifier-generated rationales rather than d…

Cited by 0SourceScholar
2026

SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation

RSS 2026poster

Robotic navigation in human environments requires a spatio-temporal semantic representation that can reconcile open-vocabulary perception with long-term environmental changes. While foundation models provide strong zero-shot recognition, their predictions are intermittent and view-dependent, and nai…

Cited by 0SourceScholar
2026

Towards 3D Object-Centric Feature Learning for Semantic Scene Completion

AAAI 2026technical

Vision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire scene, they often overlook fine-grained object-level details,

Cited by 0SourcePDFScholar
2025

ALVO: Adaptive Learning with Velocity Obstacles for UGV Navigation in Dynamic Scenes

IROS 2025

Autonomous navigation of unmanned ground vehicles (UGVs) in dynamic scenes is a challenging task that requires them to avoid obstacles and move toward the goal simultaneously. This paper proposes ALVO, an adaptive learning policy that leverages velocity obstacles for UGV navigation. ALVO employs an

Cited by 1SourceScholar
2025

Behavior Cloning Assisted Reinforcement Learning for Cable-Driven Continuum Space Robots in Sparse Reward Environments

RA-L 2025

Deep reinforcement learning (DRL) has emerged as a powerful tool for controlling cable-driven continuum space robots (CDCSRs), offering a solution that bypasses complex system modeling. However, DRL based on dense reward functions (DRLDR) requires meticulous tuning of the reward structure, whereas D

Cited by 1SourceScholar
2025

CAO-RONet: A Robust 4D Radar Odometry with Exploring More Information from Low-Quality Points

ICRA 2025

Recently, 4D millimetre-wave radar exhibits more stable perception ability than LiDAR and camera under adverse conditions (e.g. rain and fog). However, low-quality radar points hinder its application, especially the odometry task that requires a dense and accurate matching. To fully explore the pote

Cited by 3SourcecodeScholar
2025

DPR-Splat: Depth and Pose Refinement with Sparse-View 3D Gaussian Splatting for Novel View Synthesis

IROS 2025

Recent advances in 3D Gaussian Splatting have demonstrated impressive performance in novel view synthesis, particularly with dense image sets. However, its performance degrades significantly in sparse-view scenarios, primarily due to the challenge of obtaining accurate camera poses. Also, achieving

Cited by 1SourcecodeScholar
2025

DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation

ICCV 2025poster

Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-scene optimization. A key challenge lies in finding an efficient and generalizable geometric representation that seamlessly…

2025

KiGRAS: Kinematic-Driven Generative Model for Realistic Agent Simulation

RA-L 2025

Trajectory generation is a pivotal task in autonomous driving. Recent studies have introduced the autoregressive paradigm, leveraging the state transition model to approximate future trajectory distributions. This paradigm closely mirrors the real-world trajectory generation process and has achieved

Cited by 23SourceScholar
2025

LOMA: Language-assisted Semantic Occupancy Network via Triplane Mamba

AAAI 2025technical

Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information through the attention mechanism, enabling the 3D semantic o…

Cited by 1SourcePDFScholar
2025

Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis

ICASSP 2025accepted

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal synchronization adapter, our method significantly improves semantic in…

Cited by 0SourceScholar
2025

StreamMOS: Streaming Moving Object Segmentation With Multi-View Perception and Dual-Span Memory

RA-L 2025

Moving object segmentation based on LiDAR is a crucial and challenging task for autonomous driving and mobile robotics. Most approaches explore spatio-temporal information from LiDAR sequences to predict moving objects in the current frame. However, they often focus on transferring temporal cues in

Cited by 5SourcecodeScholar
2025

Target-Aware Viewpoint Generation for Active Robotic Exploration in Unknown Environments

ICRA 2025

When entering an unfamiliar environment, animals usually sweep off their surroundings to identify points of interest. In search and rescue robotics, autonomous exploration requires both coarse mapping of unknown areas and detailed target detection, which poses a significant challenge in balancing th

Cited by 0SourceScholar
2024

Dense Connector for MLLMs

NeurIPS 2024poster

*Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)?* The recent outstanding performance of MLLMs in multimodal understanding has garnered broad attention from both academia and industry. In the current MLLM rat race, the focus seems to be predominantly…

2024

Discover and Mitigate Multiple Biased Subgroups in Image Classifiers

CVPR 2024poster

Machine learning models can perform well on in-distribution data but often fail on biased subgroups that are underrepresented in the training data hindering the robustness of models for reliable applications. Such subgroups are typically unknown due to the absence of subgroup labels. Discovering bia…

2024

FairRAG: Fair Human Generation via Fair Retrieval Augmentation

CVPR 2024poster

Existing text-to-image generative models reflect or even amplify societal biases ingrained in their training data. This is especially concerning for human image generation where models are biased against certain demographic groups. Existing attempts to rectify this issue are hindered by the inherent…

Cited by 21SourcePDFScholar
2024

Learning Dual-Level Deformable Implicit Representation for Real-World Scale Arbitrary Super-Resolution

ECCV 2024poster

"Scale arbitrary super-resolution based on implicit image function gains increasing popularity since it can better represent the visual world in a continuous manner. However, existing scale arbitrary works are trained and evaluated on simulated datasets, where low-resolution images are generated fro…

2024

Mixing Left and Right-Hand Driving Data in a Hierarchical Framework With LLM Generation

RA-L 2024

Data-driven trajectory prediction is critical in autonomous vehicles, which requires high-quality data. However, discussions about the compatibility of data collected from different countries remain limited, with a typical issue being the different driving rules in various countries. Therefore, we p

Cited by 5SourceScholar
2024

Multiagent Multitraversal Multimodal Self-Driving: Open MARS Dataset

CVPR 2024poster

Large-scale datasets have fueled recent advancements in AI-based autonomous vehicle research. However these datasets are usually collected from a single vehicle's one-time pass of a certain location lacking multiagent interactions or repeated traversals of the same place. Such information could lead…

2024

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

IROS 2024

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details

Cited by 90SourcecodeScholar
2024

SeqTrack3D: Exploring Sequence Information for Robust 3D Point Cloud Tracking

ICRA 2024poster

3D single object tracking (SOT) is an important and challenging task for the autonomous driving and mobile robotics. Most existing methods perform tracking between two consecutive frames while ignoring the motion patterns of the target over a series of frames, which would cause performance degradati…

Cited by 1SourcecodeScholar
2023

A Whac-a-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others

CVPR 2023poster

Machine learning models have been found to learn shortcuts---unintended decision rules that are unable to generalize---undermining models' reliability. Previous works address this problem under the tenuous assumption that only a single shortcut exists in the training data. Real-world images are rife…

2023

MFA: Multi-layer Feature-aware Attack for Object Detection

UAI 2023poster

Physical adversarial attacks can mislead detectors in real-world scenarios and have attracted increasing attention. However, most existing works manipulate the detector’s final outputs as attack targets while ignoring the inherent characteristics of objects. This can result in attacks being trapped…

2023

Skip-Plan: Procedure Planning in Instructional Videos via Condensed Action Space Learning

ICCV 2023poster

In this paper, we propose Skip-Plan, a condensed action space learning method for procedure planning in instructional videos. Current procedure planning methods all stick to the state-action pair prediction at every timestep and generate actions adjacently. Although it coincides with human intuition…

Cited by 12PDFScholar
2022

Discover and Mitigate Unknown Biases with Debiasing Alternate Networks

ECCV 2022poster

"Deep image classifiers have been found to learn biases from datasets. To mitigate the biases, most previous methods require labels of protected attributes (e.g., age, skin tone) as full-supervision, which has two limitations: 1) it is infeasible when the labels are unavailable; 2) they are incapabl…

2022

Exploiting More Information in Sparse Point Cloud for 3D Single Object Tracking

RA-L 2022

3D single object tracking is a key task in 3D computer vision. However, the sparsity of point clouds makes it difficult to compute the similarity and locate the object, posing big challenges to the 3D tracker. Previous works tried to solve the problem and improved the tracking performance in some co

Cited by 28SourcecodeScholar
2022

StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis

CVPR 2022poster

Although progress has been made for text-to-image synthesis, previous methods fall short of generalizing to unseen or underrepresented attribute compositions in the input text. Lacking compositionality could have severe implications for robustness and fairness, e.g., inability to synthesize the face…

Cited by 61PDFcodeScholar
2021

UAV-Human: A Large Benchmark for Human Behavior Understanding With Unmanned Aerial Vehicles

CVPR 2021poster

Human behavior understanding with unmanned aerial vehicles (UAVs) is of great significance for a wide range of applications, which simultaneously brings an urgent demand of large, challenging, and comprehensive benchmarks for the development and evaluation of UAV-based models. However, existing benc…

Cited by 268PDFcodeScholar
2020

Learning a Weakly-Supervised Video Actor-Action Segmentation Model With a Wise Selection

CVPR 2020oral

We address weakly-supervised video actor-action segmentation (VAAS), which extends general video object segmentation (VOS) to additionally consider action labels of the actors. The most successful methods on VOS synthesize a pool of pseudo-annotations (PAs) and then refine them iteratively. However,…

Cited by 19PDFScholar