← Search

Bodo Rosenhahn

34 accepted papers

2026

BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship Detection

CVPR 2026

We propose Bijective Universal Scene-Specific Anomalous Relationship Detection (BUSSARD), a normalizing flow-based model for detecting anomalous relations in scene graphs, generated from images. Our work follows a multimodal approach, embedding object and relationship tokens from scene graphs with a

Cited by 0SourcecodeScholar
2026

Grouping Nodes with known Value Differences: A lossless UCT-based Abstraction Algorithm

ICLR 2026poster

A core challenge of Monte Carlo Tree Search (MCTS) is its sample efficiency, which can be addressed by building and using state and/or state-action pair abstractions in parallel to the tree search, such that information can be shared among nodes of the same layer. On the Go Abstractions in Upper Con…

Cited by 0SourcecodeScholar
2025

CHiQPM: Calibrated Hierarchical Interpretable Image Classification

NeurIPS 2025poster

Globally interpretable models are a promising approach for trustworthy AI in safety-critical domains. Alongside global explanations, detailed local explanations are a crucial complement to effectively support human experts during inference. This work proposes the Calibrated Hierarchical QPM (CHiQPM)…

Cited by 0SourceScholar
2025

QPM: Discrete Optimization for Globally Interpretable Image Classification

ICLR 2025poster

Understanding the classifications of deep neural networks, e.g. used in safety-critical situations, is becoming increasingly important. While recent models can locally explain a single decision, to provide a faithful global explanation about an accurate model’s general behavior is a more challenging…

2025

UncertainSAM: Fast and Efficient Uncertainty Quantification of the Segment Anything Model

ICML 2025poster

The introduction of the Segment Anything Model (SAM) has paved the way for numerous semantic segmentation applications. For several tasks, quantifying the uncertainty of SAM is of particular interest. However, the ambiguous nature of the class-agnostic foundation model SAM challenges current uncerta…

2024

FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

ICLR 2024poster

Text-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts. A major challenge in this task is to ensure that all frames in the edited video are visually consistent. Most recent works apply advanced text-to-image diffusion models to this task by inflati…

Cited by 74SourcePDFScholar
2024

Indoor Scene Change Understanding (SCU): Segment, Describe, and Revert Any Change

IROS 2024poster

Understanding of scene changes is crucial for embodied AI applications, such as visual room rearrangement, where the agent must revert changes by restoring the objects to their original locations or states. Visual changes between two scenes, pre- and post-rearrangement, encompass two tasks: scene ch…

Cited by 2SourceScholar
2024

Mastering Zero-Shot Interactions in Cooperative and Competitive Simultaneous Games

ICML 2024poster

The combination of self-play and planning has achieved great successes in sequential games, for instance in Chess and Go. However, adapting algorithms such as AlphaZero to simultaneous games poses a new challenge. In these games, missing information about concurrent actions of other agents is a limi…

Cited by 0SourcePDFScholar
2024

PARSAC: Accelerating Robust Multi-Model Fitting with Parallel Sample Consensus

AAAI 2024technical

We present a real-time method for robust estimation of multiple instances of geometric models from noisy data. Geometric models such as vanishing points, planar homographies or fundamental matrices are essential for 3D scene analysis. Previous approaches discover distinct model instances in an itera…

2023

Deep Reinforcement Learning for Autonomous Driving using High-Level Heterogeneous Graph Representations

ICRA 2023poster

Graph networks have recently been used for decision making in automated driving tasks for their ability to capture a variable number of traffic participants. Current high-level graph-based approaches, however, do not model the entire road network and thus must rely on handcrafted features for vehicl…

Cited by 14SourceScholar
2022

ChimeraMix: Image Classification on Small Datasets via Masked Feature Mixing

IJCAI 2022poster

Deep convolutional neural networks require large amounts of labeled data samples. For many real-world applications, this is a major limitation which is commonly treated by augmentation methods. In this work, we address the problem of learning deep neural networks on small datasets. Our proposed arch…

2022

LMGP: Lifted Multicut Meets Geometry Projections for Multi-Camera Multi-Object Tracking

CVPR 2022poster

Multi-Camera Multi-Object Tracking is currently drawing attention in the computer vision field due to its superior performance in real-world applications such as video surveillance with crowded scenes or in wide spaces. In this work, we propose a mathematically elegant multi-camera multiple object t…

Cited by 44PDFcodeScholar
2021

CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the Wild

CVPR 2021poster

Human pose estimation from single images is a challenging problem in computer vision that requires large amounts of labeled training data to be solved accurately. Unfortunately, for many human activities (e.g. outdoor sports) such training data does not exist and is hard or even impossible to acquir…

Cited by 143PDFcodeScholar
2021

Context-Aware Layout to Image Generation With Enhanced Object Appearance

CVPR 2021poster

A layout to image (L2I) generation model aims to generate a complicated image containing multiple objects (things) against natural background (stuff), conditioned on a given layout. Built upon the recent advances in generative adversarial networks (GANs), recent L2I models have made great progress.…

Cited by 65PDFcodeScholar
2021

Cuboids Revisited: Learning Robust 3D Shape Fitting to Single RGB Images

CVPR 2021poster

Humans perceive and construct the surrounding world as an arrangement of simple parametric models. In particular, man-made environments commonly consist of volumetric primitives such as cuboids or cylinders. Inferring these primitives is an important step to attain high-level, abstract scene descrip…

Cited by 32PDFcodeScholar
2021

Exploring Dynamic Context for Multi-path Trajectory Prediction

ICRA 2021poster

To accurately predict future positions of different agents in traffic scenarios is crucial for safely deploying intelligent autonomous systems in the real-world environment. However, it remains a challenge due to the behavior of a target agent being affected by other agents dynamically and there bei…

Cited by 49SourcecodeScholar
2021

Making Higher Order MOT Scalable: An Efficient Approximate Solver for Lifted Disjoint Paths

ICCV 2021poster

We present an efficient approximate message passing solver for the lifted disjoint paths problem (LDP), a natural but NP-hard model for multiple object tracking (MOT). Our tracker scales to very large instances that come from long and crowded MOT sequences. Our approximate solver enables us to proce…

Cited by 45PDFcodeScholar
2021

Probabilistic Monocular 3D Human Pose Estimation With Normalizing Flows

ICCV 2021poster

3D human pose estimation from monocular images is a highly ill-posed problem due to depth ambiguities and occlusions. Nonetheless, most existing works ignore these ambiguities and only estimate a single solution. In contrast, we generate a diverse set of hypotheses that represents the full posterior…

Cited by 145PDFcodeScholar
2021

Spatial-Temporal Transformer for Dynamic Scene Graph Generation

ICCV 2021poster

Dynamic scene graph generation aims at generating a scene graph of the given video. Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects and the temporal dependencies between frames allowing for a richer semantic inte…

Cited by 175PDFcodeScholar
2020

CONSAC: Robust Multi-Model Fitting by Conditional Sample Consensus

CVPR 2020poster

We present a robust estimator for fitting multiple parametric models of the same form to noisy measurements. Applications include finding multiple vanishing points in man-made scenes, fitting planes to architectural imagery, or estimating multiple rigid motions within the same sequence. In contrast…

Cited by 73PDFcodeScholar
2020

Lifted Disjoint Paths with Application in Multiple Object Tracking

ICML 2020poster

We present an extension to the disjoint paths problem in which additional lifted edges are introduced to provide path connectivity priors. We call the resulting optimization problem the lifted disjoint paths problem. We show that this problem is NP-hard by reduction from integer multicommodity flow…

2020

NODIS: Neural Ordinary Differential Scene Understanding

ECCV 2020poster

Semantic image understanding is a challenging topic in computer vision. It requires to detect all objects in an image, but also to identify all the relations between them. Detected objects, their labels and the discovered relations can be used to construct a scene graph which provides an abstract se…

2019

RepNet: Weakly Supervised Training of an Adversarial Reprojection Network for 3D Human Pose Estimation

CVPR 2019poster

This paper addresses the problem of 3D human pose estimation from single images. While for a long time human skeletons were parameterized and fitted to the observation by satisfying a reprojection error, nowadays researchers directly use neural networks to infer the 3D pose from the observations. Ho…

Cited by 303PDFScholar
2018

Recovering Accurate 3D Human Pose in The Wild Using IMUs and a Moving Camera

ECCV 2018poster

In this work, we propose a method that combines a single hand-held camera and a set of Inertial Measurement Units (IMUs) attached at the body limbs to estimate accurate 3D poses in the wild. This poses many new challenges: the moving camera, heading drift, cluttered background, occlusions and many p…

Cited by 1257SourcePDFScholar
2016

PSyCo: Manifold Span Reduction for Super Resolution

CVPR 2016poster

The main challenge in Super Resolution (SR) is to discover the mapping between the low- and high-resolution manifolds of image patches, a complex ill-posed problem which has recently been addressed through piecewise linear regression with promising results. In this paper we present a novel regressio…

Cited by 63PDFScholar
2015

Expanding Object Detector's Horizon: Incremental Learning Framework for Object Detection in Videos

CVPR 2015poster

Over the last several years it has been shown that image-based object detectors are sensitive to the training data and often fail to generalize to examples that fall outside the original training sample domain (e.g., videos). A number of domain adaptation (DA) techniques have been proposed to add…

Cited by 58SourcePDFScholar