← Search

Haozhi Cao

14 accepted papers

2026

SplatSSC: Decoupled Depth-Guided Gaussian Splatting for Semantic Scene Completion

AAAI 2026technical

Monocular 3D Semantic Scene Completion (SSC) is a challenging yet promising task that aims to infer dense geometric and semantic descriptions of a scene from a single image. While recent object-centric paradigms significantly improve efficiency by leveraging flexible 3D Gaussian primitives, they sti

Cited by 0SourcePDFScholar
2026

TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion

CVPR 2026

Embodied 3D Semantic Scene Completion (SSC) infers dense geometry and semantics from continuous egocentric observations. Most existing Gaussian-based methods rely on random initialization of many primitives within predefined spatial bounds, resulting in redundancy and poor scalability to unbounded s

Cited by 0SourcecodeScholar
2026

When Robots Should Say ''I Don't Know'': Benchmarking Abstention in Embodied Question Answering

CVPR 2026

Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks assume that every question must be answered, but embodied agents should know when they do not have sufficient information t

Cited by 0SourceScholar
2025

Enhancing Scene Coordinate Regression With Efficient Keypoint Detection and Sequential Information

RA-L 2025

Scene Coordinate Regression (SCR) is a visual localization technique that utilizes deep neural networks (DNN) to directly regress 2D-3D correspondences for camera pose estimation. However, current SCR methods often face challenges in handling repetitive textures and meaningless areas due to their re

Cited by 3SourcecodeScholar
2025

GERA: Geometric Embedding for Efficient Point Registration Analysis

ICRA 2025

Point cloud registration aims to provide estimated transformations to align point clouds, which plays a crucial role in pose estimation of various navigation systems, such as surgical guidance systems and autonomous vehicles. Despite the impressive performance of recent models on benchmark datasets,

Cited by 3SourceScholar
2024

MCD: Diverse Large-Scale Multi-Campus Dataset for Robot Perception

CVPR 2024highlight

Perception plays a crucial role in various robot applications. However existing well-annotated datasets are biased towards autonomous driving scenarios while unlabelled SLAM datasets are quickly over-fitted and often lack environment and domain variations. To expand the frontier of these fields we i…

Cited by 36SourcePDFScholar
2024

MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic Segmentation

ICRA 2024poster

Multi-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can achieve overall improvement, they suffer from significant class-i…

Cited by 19SourcecodeScholar
2024

Outram: One-shot Global Localization via Triangulated Scene Graph and Global Outlier Pruning

ICRA 2024poster

One-shot LiDAR localization refers to the ability to estimate the robot pose from one single point cloud, which yields significant advantages in initialization and relocalization processes. In the point cloud domain, the topic has been extensively studied as a global descriptor retrieval (i.e., loop…

Cited by 19SourcecodeScholar
2024

Reliable Spatial-Temporal Voxels For Multi-Modal Test-Time Adaptation

ECCV 2024poster

"Multi-modal test-time adaptation (MM-TTA) is proposed to adapt models to an unlabeled target domain by leveraging the complementary multi-modal inputs in an online manner. Previous MM-TTA methods for 3D segmentation rely on predictions of cross-modal information in each input frame, while they igno…

2024

SGBA: Semantic Gaussian Mixture Model-Based LiDAR Bundle Adjustment

RA-L 2024

LiDAR bundle adjustment (BA) is an effective approach to reduce the drifts in pose estimation from the front-end. Existing works on LiDAR BA usually rely on predefined geometric features for landmark representation. This reliance restricts generalizability, as the system will inevitably deteriorate

Cited by 8SourceScholar
2023

Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation

ICCV 2023poster

Continual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmenta…

Cited by 20PDFScholar
2023

Segregator: Global Point Cloud Registration with Semantic and Geometric Cues

ICRA 2023poster

This paper presents Segregator, a global point cloud registration framework that exploits both semantic information and geometric distribution to efficiently build up outlier-robust correspondences and search for inliers. Current state-of-the-art algorithms rely on point features to set up putative…

Cited by 28SourcecodeScholar
2022

Source-Free Video Domain Adaptation by Learning Temporal Consistency for Action Recognition

ECCV 2022poster

"Video-based Unsupervised Domain Adaptation (VUDA) methods improve the robustness of video models, enabling them to be applied to action recognition tasks across different environments. However, these methods require constant access to source data during the adaptation process. Yet in many real-worl…

2021

Partial Video Domain Adaptation With Partial Adversarial Temporal Attentive Network

ICCV 2021poster

Partial Domain Adaptation (PDA) is a practical and general domain adaptation scenario, which relaxes the fully shared label space assumption such that the source label space subsumes the target one. The key challenge of PDA is the issue of negative transfer caused by source-only classes. For videos,…

Cited by 36PDFcodeScholar