← Search

Holger Caesar

21 accepted papers

2026

AsyncBEV: Cross-modal flow alignment in Asynchronous 3D Object Detection

ICLR 2026poster

In autonomous driving, multi-modal perception tasks like 3D object detection typically rely on well-synchronized sensors, both at training and inference. However, despite the use of hardware- or software-based synchronization algorithms, perfect synchrony is rarely guaranteed: Sensors may operate at…

Cited by 0SourcecodeScholar
2026

GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion

ICRA 2026poster

Robust and accurate perception of dynamic objects and map elements is crucial for autonomous vehicles performing safe navigation in complex traffic scenarios. While vision-only methods have become the de facto standard due to their technical advances, they can benefit from effective and cost-efficie…

2026

ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation

CVPR 2026

Recent progress in self- and weakly supervised occupancy estimation has largely relied on 2D projection or rendering-based supervision, which suffers from geometric inconsistencies and severe depth bleeding.We thus introduce ShelfOcc, a vision-only method that overcomes these limitations without rel

Cited by 3SourceScholar
2026

TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR

CVPR 2026

LiDAR perception is fundamental to robotics, enabling machines to understand their environment in 3D. A crucial task for LiDAR-based scene understanding and navigation is ground segmentation. However, existing methods are either handcrafted for specific sensor configurations or rely on costly per-po

Cited by 0SourcecodeScholar
2025

Advancing High-Resolution and Efficient Automotive Radar Imaging through Domain-Informed 1D Deep Learning

ICASSP 2025accepted

Millimeter-wave (mmWave) radars are critical for autonomous vehicles’ perception tasks, offering reliable performance in adverse weather conditions. However, their application is often hindered by insufficient spatial resolution for detailed semantic scene interpretation. Traditional super-resolutio…

Cited by 0SourceScholar
2025

CoDa-4DGS: Dynamic Gaussian Splatting with Context and Deformation Awareness for Autonomous Driving

ICCV 2025poster

Dynamic scene rendering opens new avenues in autonomous driving by enabling closed-loop simulations with photorealistic data, which is crucial for validating end-to-end algorithms. However, the complex and highly dynamic nature of traffic environments presents significant challenges in accurately re…

Cited by 0SourcePDFScholar
2025

VoteFlow: Enforcing Local Rigidity in Self-Supervised Scene Flow

CVPR 2025poster

Scene flow estimation aims to recover per-point motion from two adjacent LiDAR scans. However, in real-world applications such as autonomous driving, points rarely move independently of others, especially for nearby points belonging to the same object, which often share the same motion. Incorporatin…

2024

BaSAL: Size-Balanced Warm Start Active Learning for LiDAR Semantic Segmentation

ICRA 2024poster

Active learning strives to reduce the need for costly data annotation, by repeatedly querying an annotator to label the most informative samples from a pool of unlabeled data, and then training a model from these samples. We identify two problems with existing active learning methods for LiDAR seman…

Cited by 5SourcecodeScholar
2024

NeuroNCAP: Photorealistic Closed-loop Safety Testing for Autonomous Driving

ECCV 2024poster

"We present a versatile NeRF-based simulator for testing autonomous driving (AD) software systems, designed with a focus on sensor-realistic closed-loop evaluation and the creation of safety-critical scenarios. The simulator learns from sequences of real-world driving sensor data and enables reconfi…

2024

OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models

ECCV 2024poster

"Panoptic Scene Graph Generation (PSG) aims to segment objects and recognize their relations, enabling the structured understanding of an image. Previous methods focus on predicting predefined object and relation categories, hence limiting their applications in the open world scenarios. With the rap…

2024

Towards learning-based planning: The nuPlan benchmark for real-world autonomous driving

ICRA 2024poster

Machine Learning (ML) has replaced handcrafted methods for perception and prediction in autonomous vehicles. Yet for the equally important planning task, the adoption of ML-based techniques is slow. We present nuPlan, the world’s first real-world autonomous driving dataset and benchmark. The benchma…

Cited by 35SourceScholar
2024

UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes

NeurIPS 2024poster

Unsupervised 3D object detection methods have emerged to leverage vast amounts of data without requiring manual labels for training. Recent approaches rely on dynamic objects for learning to detect mobile objects but penalize the detections of static instances during training. Multiple rounds of (se…

2023

HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph Generation

ICCV 2023poster

Panoptic Scene Graph generation (PSG) is a recently proposed task in image scene understanding that aims to segment the image and extract triplets of subjects, objects and their relations to build a scene graph. This task is particularly challenging for two reasons. First, it suffers from a long-tai…

Cited by 23PDFcodeScholar
2023

SliceMatch: Geometry-Guided Aggregation for Cross-View Pose Estimation

CVPR 2023poster

This work addresses cross-view camera pose estimation, i.e., determining the 3-Degrees-of-Freedom camera pose of a given ground-level image w.r.t. an aerial image of the local area. We propose SliceMatch, which consists of ground and aerial feature extractors, feature aggregators, and a pose predict…

2022

Panoptic Nuscenes: A Large-Scale Benchmark for LiDAR Panoptic Segmentation and Tracking

RA-L 2022

Panoptic scene understanding and tracking of dynamic agents are essential for robots and automated vehicles to navigate in urban environments. As LiDARs provide accurate illumination-independent geometric depictions of the scene, performing these tasks using LiDAR point clouds provides reliable pred

Cited by 243SourceScholar
2020

nuScenes: A Multimodal Dataset for Autonomous Driving

CVPR 2020poster

Robust detection and tracking of objects is crucial for the deployment of autonomous vehicle technology. Image based benchmark datasets have driven development in computer vision tasks such as object detection, tracking and segmentation of agents in the environment. Most autonomous vehicles, however…

Cited by 7379PDFcodeScholar
2019

PointPillars: Fast Encoders for Object Detection From Point Clouds

CVPR 2019poster

Object detection in point clouds is an important aspect of many robotics applications such as autonomous driving. In this paper, we consider the problem of encoding a point cloud into a format appropriate for a downstream detection pipeline. Recent literature suggests two types of encoders; fixed en…

Cited by 4630PDFcodeScholar