← Search

Simon Doll

5 accepted papers

2026

EMPERROR: A Flexible Generative Perception Error Model for Probing Self-Driving Planners

ICRA 2026poster

To handle the complexities of real-world traffic, learning planners for self-driving from data is a promising direction. While recent approaches have shown great progress, they typically assume a setting in which the ground-truth world state is available as input. However, when deployed, planning ne…

2025

Emperror: A Flexible Generative Perception Error Model for Probing Self-Driving Planners

RA-L 2025

To handle the complexities of real-world traffic, learning planners for self-driving from data is a promising direction. While recent approaches have shown great progress, they typically assume a setting in which the ground-truth world state is available as input. However, when deployed, planning ne

Cited by 2SourceScholar
2024

DualAD: Disentangling the Dynamic and Static World for End-to-End Driving

CVPR 2024poster

State-of-the-art approaches for autonomous driving integrate multiple sub-tasks of the overall driving task into a single pipeline that can be trained in an end-to-end fashion by passing latent representations between the different modules. In contrast to previous approaches that rely on a unified g…

Cited by 5SourcePDFScholar
2024

S.T.A.R.-Track: Latent Motion Models for End-to-End 3D Object Tracking With Adaptive Spatio-Temporal Appearance Representations

RA-L 2024

Following the tracking-by-attention paradigm, this letter introduces an object-centric, transformer-based framework for tracking in 3D. Traditional model-based tracking approaches incorporate the geometric effect of object- and ego motion between frames with a geometric motion model. Inspired by thi

Cited by 13SourceScholar
2022

SpatialDETR: Robust Scalable Transformer-Based 3D Object Detection from Multi-View Camera Images with Global Cross-Sensor Attention

ECCV 2022poster

"Based on the key idea of DETR this paper introduces an object-centric 3D object detection framework that operates on a limited number of 3D object queries instead of dense bounding box proposals followed by non-maximum suppression. After image feature extraction a decoder-only transformer architect…