← Search

Marius Cordts

10 accepted papers

2026

EMPERROR: A Flexible Generative Perception Error Model for Probing Self-Driving Planners

ICRA 2026poster

To handle the complexities of real-world traffic, learning planners for self-driving from data is a promising direction. While recent approaches have shown great progress, they typically assume a setting in which the ground-truth world state is available as input. However, when deployed, planning ne…

2026

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving

CVPR 2026

Retrieving rare and safety-critical driving scenarios from large-scale datasets is essential for building robust autonomous driving (AD) systems. As dataset sizes continue to grow, the key challenge shifts from collecting more data to efficiently identifying the most relevant samples. We introduce S

Cited by 0SourcecodeScholar
2025

AGO: Adaptive Grounding for Open World 3D Occupancy Prediction

ICCV 2025poster

Open-world 3D semantic occupancy prediction aims to generate a voxelized 3D representation from sensor inputs while recognizing both known and unknown objects. Transferring open-vocabulary knowledge from vision-language models (VLMs) offers a promising direction but remains challenging. However, met…

2025

Emperror: A Flexible Generative Perception Error Model for Probing Self-Driving Planners

RA-L 2025

To handle the complexities of real-world traffic, learning planners for self-driving from data is a promising direction. While recent approaches have shown great progress, they typically assume a setting in which the ground-truth world state is available as input. However, when deployed, planning ne

Cited by 2SourceScholar
2024

ADA-Track: End-to-End Multi-Camera 3D Multi-Object Tracking with Alternating Detection and Association

CVPR 2024poster

Many query-based approaches for 3D Multi-Object Tracking (MOT) adopt the tracking-by-attention paradigm utilizing track queries for identity-consistent detection and object queries for identity-agnostic track spawning. Tracking-by-attention however entangles detection and tracking queries in one emb…

2024

DualAD: Disentangling the Dynamic and Static World for End-to-End Driving

CVPR 2024poster

State-of-the-art approaches for autonomous driving integrate multiple sub-tasks of the overall driving task into a single pipeline that can be trained in an end-to-end fashion by passing latent representations between the different modules. In contrast to previous approaches that rely on a unified g…

Cited by 5SourcePDFScholar
2023

3DMOTFormer: Graph Transformer for Online 3D Multi-Object Tracking

ICCV 2023poster

Tracking 3D objects accurately and consistently is crucial for autonomous vehicles, enabling more reliable downstream tasks such as trajectory prediction and motion planning. Based on the substantial progress in object detection in recent years, the tracking-by-detection paradigm has become a popula…

Cited by 35PDFcodeScholar
2023

PowerBEV: A Powerful Yet Lightweight Framework for Instance Prediction in Bird’s-Eye View

IJCAI 2023poster

Accurately perceiving instances and predicting their future motion are key tasks for autonomous vehicles, enabling them to navigate safely in complex urban traffic. While bird’s-eye view (BEV) representations are commonplace in perception for autonomous driving, their potential in a motion predictio…

2022

Structural Knowledge Distillation for Object Detection

NeurIPS 2022accept

Knowledge Distillation (KD) is a well-known training paradigm in deep neural networks where knowledge acquired by a large teacher model is transferred to a small student. KD has proven to be an effective technique to significantly improve the student's performance for various tasks including object…

Cited by 35SourcePDFScholar
2016

The Cityscapes Dataset for Semantic Urban Scene Understanding

CVPR 2016spotlight

Visual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban scene understanding, however, no current dataset adequately…

Cited by 15494PDFScholar