← Search

Ji Zhao

24 accepted papers

2026

City-Scale Lane-Level Mapping From Crowdsourced Trajectories and Satellite Imagery

RA-L 2026

Lane-level maps are increasingly preferred over Standard-Definition (SD) and High-Definition (HD) maps, offering a better trade-off among detail richness, coverage breadth, and data freshness. However, constructing city-scale lane-level maps remains time-consuming and labor-intensive. To address the

Cited by 0SourceScholar
2026

Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better

ICLR 2026poster

As Large Language Models (LLMs) achieve remarkable empirical success through scaling model and data size, pretraining has become increasingly critical yet computationally prohibitive, hindering rapid development. Despite the availability of numerous pretrained LLMs developed at significant computati…

Cited by 0SourceScholar
2025

Full-DoF Egomotion Estimation for Event Cameras Using Geometric Solvers

CVPR 2025highlight

For event cameras, current sparse geometric solvers for egomotion estimation assume that the rotational displacements are known, such as those provided by an IMU. Thus, they can only recover the translational motion parameters. Recovering full-DoF motion parameters using a sparse geometric solver is…

2025

Online Temporal Fusion for Vectorized Map Construction in Mapless Autonomous Driving

RA-L 2025

To reduce the reliance on high-definition (HD) maps, a growing trend in autonomous driving is leveraging onboard sensors to generate vectorized maps online. However, current methods are mostly constrained by processing only single-frame inputs, which hampers their robustness and effectiveness in com

Cited by 5SourceScholar
2024

Dialogue Cross-Enhanced Central Engagement Attention Model for Real-Time Engagement Estimation

IJCAI 2024poster

Real-time engagement estimation has been an important research topic in human-computer interaction in recent years. The emergence of the NOvice eXpert Interaction (NOXI) dataset, enriched with frame-wise engagement annotations, has catalyzed a surge in research efforts in this domain. Existing featu…

2024

Enhancing Vectorized Map Perception with Historical Rasterized Maps

ECCV 2024poster

"In autonomous driving, there is growing interest in end-to-end online vectorized map perception in bird’s-eye-view (BEV) space, with an expectation that it could replace traditional high-cost offline high-definition (HD) maps. However, the accuracy and robustness of these methods can be easily comp…

2022

Affine Correspondences between Multi-Camera Systems for 6DOF Relative Pose Estimation

ECCV 2022poster

"We present a novel method to compute the 6DOF relative pose of multi-camera systems using two affine correspondences (ACs). Existing solutions to the multi-camera relative pose estimation are either restricted to special cases of motion, have too high computational complexity, or require too many p…

2021

Efficient Recovery of Multi-Camera Motion from Two Affine Correspondences

ICRA 2021poster

We propose an efficient method to estimate the relative pose of a multi-camera system from a minimum of two affine correspondences (ACs). Our solution is novel as it computes the 6DOF relative pose by utilizing a first-order rotation approximation. We directly derive a single polynomial based on the…

Cited by 9SourceScholar
2021

Learning To Identify Correct 2D-2D Line Correspondences on Sphere

CVPR 2021poster

Given a set of putative 2D-2D line correspondences, we aim to identify correct matches. Existing methods exploit the geometric constraints. They are only applicable to structured scenes with orthogonality, parallelism and coplanarity. In contrast, we propose the first approach suitable for both stru…

Cited by 4PDFScholar
2021

Minimal Cases for Computing the Generalized Relative Pose Using Affine Correspondences

ICCV 2021poster

We propose three novel solvers for estimating the relative pose of a multi-camera system from affine correspondences (ACs). A new constraint is derived interpreting the relationship of ACs and the generalized camera model. Using the constraint, we demonstrate efficient solvers for two types of motio…

Cited by 17PDFScholar
2021

Structure Reconstruction Using Ray-Point-Ray Features: Representation and Camera Pose Estimation

ICRA 2021poster

Straight line features have been increasingly utilized in visual SLAM and 3D reconstruction systems. The straight lines’ parameterization, parallel constraint, and coplanar constraint are studied in many recent works. In this paper, we explore the novel intersection constraint of straight lines for…

Cited by 3SourceScholar
2020

Globally Optimal and Efficient Vanishing Point Estimation in Atlanta World

ECCV 2020poster

Atlanta world holds for the scenes composed of a vertical dominant direction and several horizontal dominant directions. Vanishing point (VP) is the intersection of the image lines projected from parallel 3D lines. In Atlanta world, given a set of image lines, we aim to cluster them by the unknown-b…

Cited by 17SourcePDFScholar
2020

Minimal Solutions for Relative Pose With a Single Affine Correspondence

CVPR 2020poster

In this paper we present four cases of minimal solutions for two-view relative pose estimation by exploiting the affine transformation between feature points and we demonstrate efficient solvers for these cases. It is shown, that under the planar motion assumption or with knowledge of a vertical dir…

Cited by 47PDFScholar
2020

Robust and Efficient Estimation of Absolute Camera Pose for Monocular Visual Odometry

ICRA 2020poster

Given a set of 3D-to-2D point correspondences corrupted by outliers, we aim to robustly estimate the absolute camera pose. Existing methods robust to outliers either fail to guarantee high robustness and efficiency simultaneously, or require an appropriate initial pose and thus lack generality. In c…

Cited by 5SourceScholar
2019

Leveraging Structural Regularity of Atlanta World for Monocular SLAM

ICRA 2019poster

A wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocul…

Cited by 48SourceScholar
2019

Line-based Absolute and Relative Camera Pose Estimation in Structured Environments

IROS 2019poster

3D lines in structured environments encode particular regularity like parallelism and orthogonality. We leverage this structural regularity to estimate the absolute and relative camera poses. We decouple the rotation and translation, and propose a novel rotation estimation method. We decompose the a…

Cited by 22SourceScholar
2019

Quasi-Globally Optimal and Efficient Vanishing Point Estimation in Manhattan World

ICCV 2019oral

The image lines projected from parallel 3D lines intersect at a common point called the vanishing point (VP). Manhattan world holds for the scenes with three orthogonal VPs. In Manhattan world, given several lines in a calibrated image, we aim at clustering them by three unknown-but-sought VPs. The…

Cited by 37PDFScholar
2018

Robust Camera Pose Estimation via Consensus on Ray Bundle and Vector Field

IROS 2018poster

Estimating the camera pose requires point correspondences. However, in practice, correspondences are inevitably corrupted by outliers, which affects the pose estimation. We propose a general and accurate outlier removal strategy for robust camera pose estimation. The proposed strategy can detect out…

Cited by 7SourceScholar
2018

Visual Homing via Guided Locality Preserving Matching

ICRA 2018poster

This study proposes a simple yet surprisingly effective feature matching approach, termed as guided locality preserving matching (GLPM), for visual homing of panoramic images. The key idea of our approach is merely to preserve the neighborhood structures of potential true matches between two panoram…

Cited by 16SourceScholar