← Search

Simone Schaub-Meyer

15 accepted papers

2026

Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?

ICML 2026poster

Deep learning has become an essential part of computer vision, with deep neural networks (DNNs) excelling in predictive performance. However, they often fall short in other critical quality dimensions, such as robustness, calibration, or fairness. While existing studies have focused on a subset of t…

Cited by 0SourceScholar
2026

MUFASA: A Multi-Layer Framework for Slot Attention

CVPR 2026

Unsupervised object-centric learning (OCL) decomposes visual scenes into distinct entities. Slot attention is a popular approach that represents individual objects as latent vectors, called slots. Current methods obtain these slot representations solely from the last layer of a pre-trained vision tr

Cited by 0SourceScholar
2026

Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities

ICRA 2026poster

Egocentric action recognition enables robots to facilitate human-robot interactions and monitor task progress. Existing methods often rely solely on RGB videos, although additional modalities, such as audio, can improve accuracy under challenging conditions. However, most multimodal approaches assum…

2026

What is Missing? Explaining Neurons Activated by Absent Concepts

ICML 2026poster

Explainable artificial intelligence (XAI) aims to provide human-interpretable insights into the behavior of deep neural networks (DNNs), typically by estimating a simplified causal structure of the model. In existing work, this causal structure often includes relationships where the presence of a co…

Cited by 0SourceScholar
2025

ART: Adaptive Relation Tuning for Generalized Relation Prediction

ICCV 2025poster

Visual relation detection (VRD) is the task of identifying the relationships between objects in a scene. VRD models trained solely on relation detection data struggle to generalize beyond the relations on which they are trained. While prompt tuning has been used to adapt vision-language models (VLMs…

2025

Boosting Omnidirectional Stereo Matching with a Pre-trained Depth Foundation Model

IROS 2025

Omnidirectional depth perception is essential for mobile robotics applications that require scene understanding across a full 360° field of view. Camera-based setups offer a cost-effective option by using stereo depth estimation to generate dense, high-resolution depth maps without relying on expens

Cited by 0SourcecodeScholar
2025

GLASS: Guided Latent Slot Diffusion for Object-Centric Learning

CVPR 2025poster

Object-centric learning aims to decompose an input image into a set of meaningful object files (slots). These latent object representations enable a variety of downstream tasks. Yet, object-centric learning struggles on real-world datasets, which contain multiple objects of complex textures and shap…

Cited by 0SourcePDFScholar
2023

Entropy-driven Unsupervised Keypoint Representation Learning in Videos

ICML 2023poster

Extracting informative representations from videos is fundamental for effectively learning various downstream tasks. We present a novel approach for unsupervised learning of meaningful representations from videos, leveraging the concept of image spatial entropy (ISE) that quantifies the per-pixel in…

2023

FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI Methods

ICCV 2023oral

The field of explainable artificial intelligence (XAI) aims to uncover the inner workings of complex deep neural models. While being crucial for safety-critical domains, XAI inherently lacks ground-truth explanations, making its automatic evaluation an unsolved problem. We address this challenge by…

Cited by 28PDFcodeScholar
2019

Neural Inter-Frame Compression for Video Coding

ICCV 2019poster

While there are many deep learning based approaches for single image compression, the field of end-to-end learned video coding has remained much less explored. Therefore, in this work we present an inter-frame compression approach for neural video coding that can seamlessly build up on different exi…

Cited by 224PDFScholar