← Search

Carsten Rother

46 accepted papers

2026

Product-Quantised Image Representation for High-Quality Image Synthesis

ICLR 2026poster

Product quantisation (PQ) is a classical method for scalable vector encoding, yet it has seen limited usage for latent representations in high-fidelity image generation. In this work, we introduce PQGAN, a quantised image autoencoder that integrates PQ into the well-known vector quantisation (VQ) fr…

Cited by 0SourceScholar
2025

Towards Optimizing Large-Scale Multi-Graph Matching in Bioimaging

CVPR 2025poster

Multi-graph matching is an important problem in computer vision. Our task comes from bioimaging, where a set of 100 3D-microscopic images of worms have to be brought into correspondence. Surprisingly, virtually all existing methods are not applicable to this large-scale, real-world problem since the…

Cited by 0SourcePDFScholar
2024

ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control Systems

ECCV 2024oral

"The field of image synthesis has made tremendous strides forward in the last years. Besides defining the desired output image with text-prompts, an intuitive approach is to additionally use spatial guidance in form of an image, such as a depth map. In state-of-the-art approaches, this guidance is r…

2024

Discrete Cycle-Consistency Based Unsupervised Deep Graph Matching

AAAI 2024technical

We contribute to the sparsely populated area of unsupervised deep graph matching with application to keypoint matching in images. Contrary to the standard supervised approach, our method does not require ground truth correspondences between keypoint pairs. Instead, it is self-supervised by enforcing…

Cited by 1SourcePDFScholar
2022

A Comparative Study of Graph Matching Algorithms in Computer Vision

ECCV 2022poster

"The graph matching optimization problem is an essential component for many tasks in computer vision, such as bringing two deformable objects in correspondence. Naturally, a wide range of applicable algorithms have been proposed in the last decades. Since a common standard benchmark has not been dev…

2022

Neural Head Avatars From Monocular RGB Videos

CVPR 2022poster

We present Neural Head Avatars, a novel neural representation that explicitly models the surface geometry and appearance of an animatable human avatar that can be used for teleconferencing in AR/VR or other applications in the movie or games industry that rely on a digital human. Our representation…

Cited by 223PDFScholar
2022

Towards Multimodal Depth Estimation From Light Fields

CVPR 2022poster

Light field applications, especially light field rendering and depth estimation, developed rapidly in recent years. While state-of-the-art light field rendering methods handle semi-transparent and reflective objects well, depth estimation methods either ignore these cases altogether or only deliver…

Cited by 14PDFScholar
2021

Fusion Moves for Graph Matching

ICCV 2021poster

We contribute to approximate algorithms for the quadratic assignment problem also known as graph matching. Inspired by the success of the fusion moves technique developed for multilabel discrete Markov random fields, we investigate its applicability to graph matching. In particular, we show how fusi…

Cited by 16PDFcodeScholar
2021

Generative Classifiers as a Basis for Trustworthy Image Classification

CVPR 2021poster

With the maturing of deep learning systems, trustworthiness is becoming increasingly important for model assessment. We understand trustworthiness as the combination of explainability and robustness. Generative classifiers (GCs) are a promising class of models that are said to naturally accomplish t…

Cited by 60PDFcodeScholar
2021

On the Limits of Pseudo Ground Truth in Visual Camera Re-Localisation

ICCV 2021poster

Benchmark datasets that measure camera pose accuracy have driven progress in visual re-localisation research. To obtain poses for thousands of images, it is common to use a reference algorithm to generate pseudo ground truth. Popular choices include Structure-from-Motion (SfM) and Simultaneous-Local…

Cited by 75PDFcodeScholar
2021

Self-Supervised Object Detection via Generative Image Synthesis

ICCV 2021poster

We present SSOD -- the first end-to-end analysis-by-synthesis framework with controllable GANs for the task of self-supervised object detection. We use collections of real-world images without bounding box annotations to learn to synthesize and detect objects. We leverage controllable GANs to synthe…

Cited by 15PDFcodeScholar
2020

A Primal-Dual Solver for Large-Scale Tracking-by-Assignment

AISTATS 2020poster

We propose a fast approximate solver for the combinatorial problem known as tracking-by-assignment, which we apply to cell tracking. The latter plays a key role in discovery in many life sciences, especially in cell and developmental biology. So far, in the most general setting this problem was addr…

2020

CONSAC: Robust Multi-Model Fitting by Conditional Sample Consensus

CVPR 2020poster

We present a robust estimator for fitting multiple parametric models of the same form to noisy measurements. Applications include finding multiple vanishing points in man-made scenes, fitting planes to architectural imagery, or estimating multiple rigid motions within the same sequence. In contrast…

Cited by 73PDFcodeScholar
2020

Disentanglement by Nonlinear ICA with General Incompressible-flow Networks (GIN)

ICLR 2020spotlight

A central question of representation learning asks under which conditions it is possible to reconstruct the true latent variables of an arbitrarily complex generative process. Recent breakthrough work by Khemakhem et al. (2019) on nonlinear ICA has answered this question for a broad class of conditi…

Cited by 151SourceScholar
2020

Increasing the Robustness of Semantic Segmentation Models with Painting-by-Numbers

ECCV 2020poster

For safety-critical applications such as autonomous driving, CNNs have to be robust with respect to unavoidable image corruptions, such as image noise. While previous works addressed the task of robust prediction in the context of full-image classification, we consider it for dense semantic segmenta…

Cited by 26SourcePDFScholar
2020

Reinforced Feature Points: Optimizing Feature Detection and Description for a High-Level Task

CVPR 2020oral

We address a core problem of computer vision: Detection and description of 2D feature points for image matching. For a long time, hand-crafted designs, like the seminal SIFT algorithm, were unsurpassed in accuracy and efficiency. Recently, learned feature detectors emerged that implement detection a…

Cited by 97PDFcodeScholar
2020

Self-Supervised Viewpoint Learning From Image Collections

CVPR 2020poster

Training deep neural networks to estimate the viewpoint of objects requires large labeled training datasets. However, manually labeling viewpoints is notoriously hard, error-prone, and time-consuming. On the other hand, it is relatively easy to mine many unlabeled images of an object category from t…

Cited by 45PDFcodeScholar
2020

Taxonomy of Dual Block-Coordinate Ascent Methods for Discrete Energy Minimization

AISTATS 2020poster

We consider the maximum-a-posteriori inference problem in discrete graphical models and study solvers based on the dual block-coordinate ascent rule. We map all existing solvers in a single framework, allowing for a better understanding of their design principles. We theoretically show that some blo…

2020

Training Normalizing Flows with the Information Bottleneck for Competitive Generative Classification

NeurIPS 2020oral

The Information Bottleneck (IB) objective uses information theory to formulate a task-performance versus robustness trade-off. It has been successfully applied in the standard discriminative classification setting. We pose the question whether the IB can also be used to train generative likelihood m…

2019

Analyzing Inverse Problems with Invertible Neural Networks

ICLR 2019poster

For many applications, in particular in natural science, the task is to determine hidden system parameters from a set of measurements. Often, the forward process from parameter- to measurement-space is well-defined, whereas the inverse problem is ambiguous: multiple parameter sets can result in the…

Cited by 687SourcePDFScholar
2018

BOP: Benchmark for 6D Object Pose Estimation

ECCV 2018poster

We propose a benchmark for 6D pose estimation of a rigid object from a single RGB-D input image. The training data consists of a texture-mapped 3D object model or images of the object in known 6D poses. The benchmark comprises of: i) eight datasets in a unified format that cover different practical…

2018

MPLP++: Fast, Parallel Dual Block-Coordinate Ascent for Dense Graphical Models

ECCV 2018poster

Dense, discrete Graphical Models with pairwise potentials are a powerful class of models which are employed in state-of-the-art computer vision and bio-imaging applications. This work introduces a new MAP-solver, based on the popular Dual Block-Coordinate Ascent principle. Surprisingly, by making a…

Cited by 22SourcePDFScholar
2018

Trust Your Model: Light Field Depth Estimation With Inline Occlusion Handling

CVPR 2018poster

We address the problem of depth estimation from light-field images. Our main contribution is a new way to handle occlusions which improves general accuracy and quality of object borders. In contrast to all prior work we work with a model which directly incorporates both depth and occlusion, using a…

Cited by 78SourcePDFScholar
2017

A Study of Lagrangean Decompositions and Dual Ascent Solvers for Graph Matching

CVPR 2017poster

We study the quadratic assignment problem, in computer vision also known as graph matching. Two leading solvers for this problem optimize the Lagrange decomposition duals with sub-gradient and dual ascent (also known as message passing) updates. We explore this direction further and propose several…

Cited by 68PDFcodeScholar
2017

Analyzing modular CNN architectures for joint depth prediction and semantic segmentation

ICRA 2017poster

This paper addresses the task of designing a modular neural network architecture that jointly solves different tasks. As an example we use the tasks of depth estimation and semantic segmentation given a single RGB image. The main focus of this work is to analyze the cross-modality influence between…

Cited by 82SourceScholar
2017

Bounding Boxes, Segmentations and Object Coordinates: How Important Is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios?

ICCV 2017poster

Existing methods for 3D scene flow estimation often fail in the presence of large displacement or local ambiguities, e.g., at texture-less or reflective surfaces. However, these challenges are omnipresent in dynamic road scenes, which is the focus of this work. Our main contribution is to overcome t…

Cited by 189PDFScholar
2017

DSAC - Differentiable RANSAC for Camera Localization

CVPR 2017oral

RANSAC is an important algorithm in robust optimization and a central building block for many computer vision applications. In recent years, traditionally hand-crafted pipelines have been replaced by deep learning pipelines, which can be trained in an end-to-end fashion. However, RANSAC has so far n…

Cited by 737PDFcodeScholar
2017

Global Hypothesis Generation for 6D Object Pose Estimation

CVPR 2017spotlight

This paper addresses the task of estimating the 6D-pose of a known 3D object from a single RGB-D image. Most modern approaches solve this task in three steps: i) compute local features; ii) generate a pool of pose-hypotheses; iii) select and refine a pose from the pool. This work focuses on the seco…

Cited by 156PDFScholar
2017

InstanceCut: From Edges to Instances With MultiCut

CVPR 2017poster

This work addresses the task of instance-aware semantic segmentation. Our key motivation is to design a simple method with a new modelling-paradigm, which therefore has a different trade-off between advantages and disadvantages compared to known approaches. Our approach, we term InstanceCut, represe…

Cited by 324PDFScholar
2017

Joint Graph Decomposition & Node Labeling: Problem, Algorithms, Applications

CVPR 2017poster

We state a combinatorial optimization problem whose feasible solutions define both a decomposition and a node labeling of a given graph. This problem offers a common mathematical abstraction of seemingly unrelated computer vision tasks, including instance-separating semantic segmentation, articulate…

Cited by 131PDFcodeScholar
2017

PoseAgent: Budget-Constrained 6D Object Pose Estimation via Reinforcement Learning

CVPR 2017poster

State-of-the-art computer vision algorithms often achieve efficiency by making discrete choices about which hypotheses to explore next. This allows allocation of computational resources to promising candidates, however, such decisions are non-differentiable. As a result, these algorithms are hard t…

Cited by 60PDFScholar
2017

Random forests versus Neural Networks — What's best for camera localization?

ICRA 2017poster

This work addresses the task of camera localization in a known 3D scene given a single input RGB image. State-of-the-art approaches accomplish this in two steps: firstly, regressing for every pixel in the image its 3D scene coordinate and subsequently, using these coordinates to estimate the final 6…

Cited by 90SourceScholar
2016

Convexity Shape Constraints for Image Segmentation

CVPR 2016poster

Segmenting an image into multiple components is a central task in computer vision. In many practical scenarios, prior knowledge about plausible components is available. Incorporating such prior knowledge into models and algorithms for image segmentation is highly desirable, yet can be non-trivial. I…

Cited by 31PDFScholar
2016

Joint M-Best-Diverse Labelings as a Parametric Submodular Minimization

NeurIPS 2016poster

We consider the problem of jointly inferring the $M$-best diverse labelings for a binary (high-order) submodular energy of a graphical model. Recently, it was shown that this problem can be solved to a global optimum, for many practically interesting diversity measures. It was noted that the labelin…

Cited by 19SourcePDFScholar
2016

Uncertainty-Driven 6D Pose Estimation of Objects and Scenes From a Single RGB Image

CVPR 2016poster

In recent years, the task of estimating the 6D pose of object instances and complete scenes, i.e. camera localization, from a single input image has received considerable attention. Consumer RGB-D cameras have made this feasible, even for difficult, texture-less objects and scenes. In this work, we…

Cited by 627PDFScholar
2015

Inferring M-Best Diverse Labelings in a Single One

ICCV 2015poster

We consider the task of finding M-best diverse solutions in a graphical model. In a previous work by Batra et al. an algorithmic approach for finding such solutions was proposed, and its usefulness was shown in numerous applications. Contrary to previous work we propose a novel formulation of the pr…

Cited by 51PDFScholar
2015

Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images

ICCV 2015poster

Analysis-by-synthesis has been a successful approach for many tasks in computer vision, such as 6D pose estimation of an object in an RGB-D image which is the topic of this work. The idea is to compare the observation with the output of a forward process, such as a rendered image of the object of in…

Cited by 263PDFScholar
2015

M-Best-Diverse Labelings for Submodular Energies and Beyond

NeurIPS 2015poster

We consider the problem of finding M best diverse solutions of energy minimization problems for graphical models. Contrary to the sequential method of Batra et al., which greedily finds one solution after another, we infer all $M$ solutions jointly. It was shown recently that such jointly inferred l…

Cited by 26SourcePDFScholar