← Search

Zhiyu Xiang

22 accepted papers

2026

SDNet: LiDAR Semantic Scene Completion with Sparse-Dense Fusion and Input-Aware Label Refinement

AAAI 2026technical

LiDAR Semantic Scene Completion (SSC) in autonomous driving requires predicting both dense occupancy and semantic labels from sparse input point cloud. Existing methods typically adopt cascaded architecture for feature dilation and semantic abstraction, which blurs distinctive geometric patterns and

Cited by 0SourcePDFScholar
2025

Privacy-Preserving V2X Collaborative Perception Integrating Unknown Collaborators

AAAI 2025technical

Vehicle-to-everything (V2X) collaborative perception has recently gained increasing attention in autonomous driving due to its ability to enhance scene understanding by integrating information from other collaborators, e.g. vehicles or infrastructure. Existing algorithms usually share deep features…

Cited by 0SourcePDFScholar
2025

RayFusion: Ray Fusion Enhanced Collaborative Visual Perception

NeurIPS 2025poster

Collaborative visual perception methods have gained widespread attention in the autonomous driving community in recent years due to their ability to address sensor limitation problems. However, the absence of explicit depth information often makes it difficult for camera-based perception systems, e.…

Cited by 0SourceScholar
2025

SCKD: Semi-Supervised Cross-Modality Knowledge Distillation for 4D Radar Object Detection

AAAI 2025technical

3D object detection is one of the fundamental perception tasks for autonomous vehicles. Fulfilling such a task with a 4D millimeter-wave radar is very attractive since the sensor is able to acquire 3D point clouds similar to Lidar while maintaining robust measurements under adverse weather. However,…

2024

Adaptive Multi-Modal Cross-Entropy Loss for Stereo Matching

CVPR 2024poster

Despite the great success of deep learning in stereo matching recovering accurate disparity maps is still challenging. Currently L1 and cross-entropy are the two most widely used losses for stereo network training. Compared with the former the latter usually performs better thanks to its probability…

2024

Exploiting Depth Priors for Few-Shot Neural Radiance Field Reconstruction

RA-L 2024

The performance of neural radiance field technologies deteriorates rapidly when sparse views are used as input. In this paper, we propose a simulated viewpoint enhancement for surface reconstruction that extracts diverse geometric features from the depth to address this limitation. We design a novel

Cited by 0SourceScholar
2024

IFTR: An Instance-Level Fusion Transformer for Visual Collaborative Perception

ECCV 2024poster

"Multi-agent collaborative perception has emerged as a widely recognized technology in the field of autonomous driving in recent years. However, current collaborative perception predominantly relies on LiDAR point clouds, with significantly less attention given to methods using camera images. This s…

2024

RISurConv: Rotation Invariant Surface Attention-Augmented Convolutions for 3D Point Cloud Classification and Segmentation

ECCV 2024oral

"Despite the progress on 3D point cloud deep learning, most prior works focus on learning features that are invariant to translation and point permutation, and very limited efforts have been devoted for rotation invariant property. Several recent studies achieve rotation invariance at the cost of lo…

2023

TransAPR: Absolute Camera Pose Regression With Spatial and Temporal Attention

RA-L 2023

Visual relocalization aims to estimate the absolute camera pose from an image or sequential images. Recent works tackle this problem by exploiting deep neural networks to regress camera poses. However, spatial and temporal clues from sequential images still remain underexplored, resulting in inaccur

Cited by 9SourceScholar
2022

CVFNet: Real-time 3D Object Detection by Learning Cross View Features

IROS 2022poster

In recent years 3D object detection from LiDAR point clouds has made great progress thanks to the development of deep learning technologies. Although voxel or point based methods are popular in 3D object detection, they usually involve time-consuming operations such as 3D convolutions on voxels or b…

Cited by 20SourceScholar
2022

Cross-Modal Knowledge Distillation for Depth Privileged Monocular Visual Odometry

RA-L 2022

Most self-supervised monocular visual odometry (VO) suffer from the scale ambiguity problem. A promising way to address this problem is to introduce additional information for training. In this work, we propose a new depth privileged framework to learn a monocular VO. It assumes that sparse depth is

Cited by 9SourceScholar
2022

Homography Loss for Monocular 3D Object Detection

CVPR 2022poster

Monocular 3D object detection is an essential task in autonomous driving. However, most current methods consider each 3D object in the scene as an independent training sample, while ignoring their inherent geometric relations, thus inevitably resulting in a lack of leveraging spatial constraints. In…

Cited by 59PDFcodeScholar
2021

Real-Time Instance Segmentation With Discriminative Orientation Maps

ICCV 2021poster

Although instance segmentation has made considerable advancement over recent years, it's still a challenge to design high accuracy algorithms with real-time performance. In this paper, we propose a real-time instance segmentation framework termed OrienMask. Upon the one-stage object detector YOLOv3,…

Cited by 22PDFcodeScholar
2020

DSSF-net: Dual-Task Segmentation and Self-supervised Fitting Network for End-to-End Lane Mark Detection

IROS 2020poster

Lane mark detection is one of the key tasks for autonomous driving systems. Accurate detection of lane marks under complex urban environments remains a challenge. In this paper, an end-to-end lane mark detection network named DSSF-net, which is capable of directly outputting the accurate fitted lane…

Cited by 1SourceScholar
2019

3D Reconstruction by Single Camera Omnidirectional Multi-Stereo System

IROS 2019poster

Omnidirectional catadioptric systems are popular in robotic applications thanks to their large field of view. For 3D scene reconstruction in a single shot, usually two different catadioptric cameras are needed. More cameras may contribute to better reconstruction while larger mounting space and high…

Cited by 0SourceScholar
2018

A Multi-Position Joint Particle Filtering Method for Vehicle Localization in Urban Area

IROS 2018poster

Robust localization is a prerequisite for autonomous vehicles. Traditional visual localization methods like visual odometry suffer error accumulation on long range navigation. In this paper, a flexible road map based probabilistic filtering method is proposed to tackle this problem. To effectively m…

Cited by 5SourceScholar