← Search

Xinhao Liu

15 accepted papers

2026

Monocular Localization With Vector HD Maps Using Geometric-Context Data Association

RA-L 2026

Due to their low cost and wide field of view, monocular cameras hold great promise for visual localization. Nevertheless, significant challenges persist in associating detected semantic features with map landmarks, arising from sensing noise, depth cues, and dynamic driving scenes. Most existing app

Cited by 0SourceScholar
2026

Thinking in 360deg: Humanoid Visual Search in the Wild

CVPR 2026

Humans rely on the synergistic control of head (cephalomotor) and eye (oculomotor) to efficiently search for visual information in 360deg. However, prior approaches to visual search are limited to a static image, neglecting the physical embodiment and its interaction with the 3D world. How can we de

Cited by 0SourcecodeScholar
2026

Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI

CVPR 2026

Reproducible closed-loop evaluation remains a major bottleneck in Embodied AI such as visual navigation. A promising path forward is high-fidelity simulation that combines photorealistic sensor rendering with geometrically grounded interaction in complex, open-world urban environments. Although rece

Cited by 0SourcecodeScholar
2025

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

CVPR 2025poster

Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods struggle in map-free or off-street settings, limiting the deployment of autonomous…

2025

Equivariant Masked Position Prediction for Efficient Molecular Representation

ICLR 2025poster

Graph neural networks (GNNs) have shown considerable promise in computational chemistry. However, the limited availability of molecular data raises concerns regarding GNNs' ability to effectively capture the fundamental principles of physics and chemistry, which constrains their generalization capab…

2025

Hierarchical Refinement: Optimal Transport to Infinity and Beyond

ICML 2025oral

Optimal transport (OT) has enjoyed great success in machine learning as a principled way to align datasets via a least-cost correspondence, driven in large part by the runtime efficiency of the Sinkhorn algorithm (Cuturi, 2013). However, Sinkhorn has quadratic space complexity in the number of point…

Cited by 0SourcePDFScholar
2025

Self-Supervised Place Recognition by Refining Temporal and Featural Pseudo Labels From Panoramic Data

RA-L 2025

Visual place recognition (VPR) using deep networks has achieved state-of-the-art performance. However, most of them require a training set with ground truth sensor poses to obtain positive and negative samples of each observation's spatial neighborhood for supervised learning. When such information

Cited by 6SourcecodeScholar
2024

LiDAR-based 4D Occupancy Completion and Forecasting

IROS 2024poster

Scene completion and forecasting are two popular perception problems in research for mobile agents like autonomous vehicles. Existing approaches treat the two problems in isolation, resulting in a separate perception of the two aspects. In this paper, we introduce a novel LiDAR perception task of Oc…

Cited by 18SourcecodeScholar
2024

Low-Rank Optimal Transport through Factor Relaxation with Latent Coupling

NeurIPS 2024poster

Optimal transport (OT) is a general framework for finding a minimum-cost transport plan, or coupling, between probability distributions, and has many applications in machine learning. A key challenge in applying OT to massive datasets is the quadratic scaling of the coupling matrix with the size of…

Cited by 4SourcePDFScholar
2024

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

IROS 2024

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details

Cited by 90SourcecodeScholar
2023

DeepMapping2: Self-Supervised Large-Scale LiDAR Map Optimization

CVPR 2023poster

LiDAR mapping is important yet challenging in self-driving and mobile robotics. To tackle such a global point cloud registration problem, DeepMapping converts the complex map estimation into a self-supervised training of simple deep networks. Despite its broad convergence range on small datasets, De…

2017

Deep salience map guided arbitrary direction scene text recognition

ICASSP 2017accepted

Irregular scene text such as curved, rotated or perspective texts commonly appear in natural scene images due to different camera view points, special design purposes etc. In this work, we propose a text salience map guided model to recognize these arbitrary direction scene texts. We train a deep Fu…

Cited by 0SourceScholar
2016

Scene text recognition with high performance CNN classifier and efficient word inference

ICASSP 2016accepted

The recognition of text in natural scene images is a practical yet challenging task due to the large variations in backgrounds, textures, fonts, and illumination conditions. In this paper, we propose a highly accurate character recognition model by utilizing the representational power of a specially…

Cited by 0SourceScholar