← Search

Stefan Roth

52 accepted papers

2026

Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?

ICML 2026poster

Deep learning has become an essential part of computer vision, with deep neural networks (DNNs) excelling in predictive performance. However, they often fall short in other critical quality dimensions, such as robustness, calibration, or fairness. While existing studies have focused on a subset of t…

Cited by 0SourceScholar
2026

Hardware-Efficient Cognitive Radar: Multi-Target Detection with RL-Driven Transmissive RIS

ICASSP 2026oral

Cognitive radar has emerged as a key paradigm for next-generation sensing, enabling adaptive, intelligent operation in dynamic and complex environments. Yet, conventional cognitive multiple-input multiple-output (MIMO) radars offer strong detection performance but suffer from high hardware complexit…

Cited by 0SourcePDFScholar
2026

INSID3: Training-Free In-Context Segmentation with DINOv3

CVPR 2026

In-context segmentation (ICS) aims to segment arbitrary concepts, e.g., objects, parts, or personalized instances, given one annotated visual examples. Existing work relies on (i) fine-tuning vision foundation models (VFMs), which improves in-domain results but harms generalization, or (ii) combines

Cited by 0SourcecodeScholar
2026

MARCO: Navigating the Unseen Space of Semantic Correspondence

CVPR 2026

Recent advances in semantic correspondence rely on dual-encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion-parameter models generalize poorly beyond training keypoints, revealing a gap between benchmark performance and real-world usability, where queried

Cited by 0SourcecodeScholar
2026

MUFASA: A Multi-Layer Framework for Slot Attention

CVPR 2026

Unsupervised object-centric learning (OCL) decomposes visual scenes into distinct entities. Slot attention is a popular approach that represents individual objects as latent vectors, called slots. Current methods obtain these slot representations solely from the last layer of a pre-trained vision tr

Cited by 0SourceScholar
2026

Scene-Centric Unsupervised Video Panoptic Segmentation

CVPR 2026

Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision. Existing unsupervised scene understanding works mainly focuse

Cited by 0SourceScholar
2026

What is Missing? Explaining Neurons Activated by Absent Concepts

ICML 2026poster

Explainable artificial intelligence (XAI) aims to provide human-interpretable insights into the behavior of deep neural networks (DNNs), typically by estimating a simplified causal structure of the model. In existing work, this causal structure often includes relationships where the presence of a co…

Cited by 0SourceScholar
2025

ART: Adaptive Relation Tuning for Generalized Relation Prediction

ICCV 2025poster

Visual relation detection (VRD) is the task of identifying the relationships between objects in a scene. VRD models trained solely on relation detection data struggle to generalize beyond the relations on which they are trained. While prompt tuning has been used to adapt vision-language models (VLMs…

2025

Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model

IROS 2025

Omnidirectional depth perception is essential for mobile robotics applications that require scene understanding across a full 360° field of view. Camera-based setups offer a cost-effective option by using stereo depth estimation to generate dense, high-resolution depth maps without relying on expens

Cited by 0SourcecodeScholar
2025

Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion

ICCV 2025poster

Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC in an unsupervised setting. Our novel method, SceneDINO, adapts techniques from…

2025

GLASS: Guided Latent Slot Diffusion for Object-Centric Learning

CVPR 2025poster

Object-centric learning aims to decompose an input image into a set of meaningful object files (slots). These latent object representations enable a variety of downstream tasks. Yet, object-centric learning struggles on real-world datasets, which contain multiple objects of complex textures and shap…

Cited by 0SourcePDFScholar
2025

Scene-Centric Unsupervised Panoptic Segmentation

CVPR 2025highlight

Unsupervised panoptic segmentation aims to partition an image into semantically meaningful regions and distinct object instances without training on manually annotated data. In contrast to prior work on unsupervised panoptic scene understanding, we eliminate the need for object-centric training data…

2023

FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI Methods

ICCV 2023oral

The field of explainable artificial intelligence (XAI) aims to uncover the inner workings of complex deep neural models. While being crucial for safety-critical domains, XAI inherently lacks ground-truth explanations, making its automatic evaluation an unsolved problem. We address this challenge by…

Cited by 28PDFcodeScholar
2023

Vision Relation Transformer for Unbiased Scene Graph Generation

ICCV 2023poster

Recent years have seen a growing interest in Scene Graph Generation (SGG), a comprehensive visual scene understanding task that aims to predict entity relationships using a relation encoder-decoder pipeline stacked on top of an object encoder-decoder backbone. Unfortunately, current SGG methods suff…

Cited by 21PDFcodeScholar
2022

xGQA: Cross-Lingual Visual Question Answering

ACL 2022findings

Recent advances in multimodal vision and language modeling have predominantly focused on the English language, mostly due to the lack of multilingual multimodal datasets to steer modeling efforts. In this work, we address this gap and provide xGQA, a new multilingual evaluation benchmark for the vis…

2020

Deep Wiener Deconvolution: Wiener Meets Deep Learning for Image Deblurring

NeurIPS 2020oral

We present a simple and effective approach for non-blind image deblurring, combining classical techniques and deep learning. In contrast to existing methods that deblur the image directly in the standard image space, we propose to perform an explicit deconvolution process in a feature space by integ…

Cited by 156SourcePDFScholar
2020

Driving Style Encoder: Situational Reward Adaptation for General-Purpose Planning in Automated Driving

ICRA 2020poster

General-purpose planning algorithms for automated driving combine mission, behavior, and local motion planning. Such planning algorithms map features of the environment and driving kinematics into complex reward functions. To achieve this, planning experts often rely on linear reward functions. The…

Cited by 12SourceScholar
2020

Normalizing Flows With Multi-Scale Autoregressive Priors

CVPR 2020poster

Flow-based generative models are an important class of exact inference models that admit efficient inference and sampling for image synthesis. Owing to the efficiency constraints on the design of the flow layers, e.g. split coupling flow layers in which approximately half the pixels do not undergo f…

Cited by 14PDFcodeScholar
2020

Planning on the fast lane: Learning to interact using attention mechanisms in path integral inverse reinforcement learning

IROS 2020poster

General-purpose trajectory planning algorithms for automated driving utilize complex reward functions to perform a combined optimization of strategic, behavioral, and kinematic features. The specification and tuning of a single reward function is a tedious task and does not generalize over a large s…

Cited by 12SourceScholar
2019

Driving with Style: Inverse Reinforcement Learning in General-Purpose Planning for Automated Driving

IROS 2019poster

Behavior and motion planning play an important role in automated driving. Traditionally, behavior planners instruct local motion planners with predefined behaviors. Due to the high scene complexity in urban environments, unpredictable situations may occur in which behavior planners fail to match pre…

Cited by 71SourceScholar
2016

The Cityscapes Dataset for Semantic Urban Scene Understanding

CVPR 2016spotlight

Visual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban scene understanding, however, no current dataset adequately…

Cited by 15494PDFScholar