← Search

Konrad Schindler

60 accepted papers

2026

Continuous Space-Time Video Super-Resolution with 3D Fourier Fields

ICLR 2026poster

We introduce a novel formulation for continuous space-time video super-resolution. Instead of decoupling the representation of a video sequence into separate spatial and temporal components and relying on brittle, explicit frame warping for motion compensation, we encode video as a continuous, spati…

Cited by 0SourcecodeScholar
2026

Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding

CVPR 2026

The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present a controllable, direct end-to-end method for upgrading a conventional video to a binocular one. Our approach, based on (conditional) latent diffusion, avoids artifacts due to explicit dep

Cited by 0SourcecodeScholar
2026

FastVMT: Eliminating Redundancy in Video Motion Transfer

ICLR 2026poster

Video motion transfer aims to synthesize videos by generating visual content according to a text prompt while transferring the motion pattern observed in a reference video. Recent methods predominantly use the Diffusion Transformer (DiT) architecture. To achieve satisfactory runtime, several methods…

Cited by 0SourceScholar
2026

FireScope: Wildfire Risk Raster Prediction With a Chain-of-Thought Oracle

CVPR 2026

Predicting wildfire risk is a reasoning-intensive spatial problem that requires the integration of visual, climatic, and geographic factors to infer continuous risk maps. Existing methods lack the causal reasoning and multimodal understanding required for reliable generalization. We introduce FireSc

Cited by 0SourcecodeScholar
2026

LitePT: Lighter Yet Stronger Point Transformer

CVPR 2026

Modern neural architectures for 3D point cloud processing contain both convolutional layers and attention blocks, but the best way to assemble them remains unclear. We analyse the role of different computational blocks in 3D point cloud networks and find an intuitive behaviour: convolution is adequa

Cited by 0SourcecodeScholar
2026

Making Foundation Models Probabilistic via Singular Value Ensembles

ICML 2026poster

Foundation models have become a dominant paradigm in machine learning, achieving remarkable performance across diverse tasks through large-scale pretraining. However, these models often yield overconfident, uncalibrated predictions. The standard approach to quantifying epistemic uncertainty, trainin…

Cited by 0SourceScholar
2026

Robust Promptable Video Object Segmentation

CVPR 2026

The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deployment in safety-critical domains. This paper offers the first comprehensive study on robust PVOS (RobustPVOS). We first construct a new, comprehensive benchm

Cited by 0SourcecodeScholar
2026

SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

ICML 2026poster

Despite recent successes, *test-time scaling* $-$i.e., dynamically expanding the token budget during inference as needed$-$ remains brittle for vision-language models (VLMs): unstructured chains-of-thought about images entangle perception and reasoning, leading to long, disorganized contexts where s…

Cited by 0SourceScholar
2026

Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator

ICLR 2026oral

The rapid progress of large, pretrained models for both visual content generation and 3D reconstruction opens up new possibilities for text-to-3D generation. Intuitively, one could obtain a formidable 3D scene generator if one were able to combine the power of a modern latent text-to-video model as…

Cited by 0SourcecodeScholar
2026

Understanding, Accelerating, and Improving MeanFlow Training

CVPR 2026

MeanFlow promises high-quality generative modeling in few steps, by jointly learning instantaneous and average velocity fields. Yet, the underlying training dynamics remain unclear. We analyze the interaction between the two velocities and find: (i) well-established instantaneous velocity is a prere

Cited by 0SourcecodeScholar
2025

A Unified Solution to Video Fusion: From Multi-Frame Learning to Benchmarking

NeurIPS 2025spotlight

The real world is dynamic, yet most image fusion methods process static frames independently, ignoring temporal correlations in videos and leading to flickering and temporal inconsistency. To address this, we propose Unified Video Fusion (UniVF), a novel and unified framework for video fusion that l…

Cited by 0SourcecodeScholar
2025

A Variational Perspective on Generative Protein Fitness Optimization

ICML 2025poster

The goal of protein fitness optimization is to discover new protein variants with enhanced fitness for a given use. The vast search space and the sparsely populated fitness landscape, along with the discrete nature of protein sequences, pose significant challenges when trying to determine the gradie…

Cited by 0SourcePDFScholar
2025

CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation

ICLR 2025spotlight

We introduce a novel method for generating 360° panoramas from text prompts or images. Our approach leverages recent advances in 3D generation by employing multi-view diffusion models to jointly synthesize the six faces of a cubemap. Unlike previous methods that rely on processing equirectangular pr…

Cited by 3SourcePDFScholar
2025

GALA: Geometry-Aware Local Adaptive Grids for Detailed 3D Generation

ICLR 2025poster

We propose GALA, a novel representation of 3D shapes that (i) excels at capturing and reproducing complex geometry and surface details, (ii) is computationally efficient, and (iii) lends itself to 3D generative modelling with modern, diffusion-based schemes. The key idea of GALA is to exploit both t…

2025

Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion

ICCV 2025poster

Depth completion upgrades sparse depth measurements into dense depth maps, guided by a conventional image. Existing methods for this highly ill-posed task operate in tightly constrained settings, and tend to struggle when applied to images outside the training domain, as well as when the available d…

2025

Solving Inverse Problems with FLAIR

NeurIPS 2025poster

Flow-based latent generative models such as Stable Diffusion 3 are able to generate images with remarkable quality, even enabling photorealistic text-to-image generation. Their impressive performance suggests that these models should also constitute powerful priors for inverse imaging problems, but…

Cited by 0SourcecodeScholar
2024

AGILE3D: Attention Guided Interactive Multi-object 3D Segmentation

ICLR 2024poster

During interactive segmentation, a model and a user work together to delineate objects of interest in a 3D point cloud. In an iterative process, the model assigns each data point to an object (or the background), while the user corrects errors in the resulting segmentation and feeds them back into t…

2024

BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation

NeurIPS 2024poster

By training over large-scale datasets, zero-shot monocular depth estimation (MDE) methods show robust performance in the wild but often suffer from insufficient detail. Although recent diffusion-based MDE approaches exhibit a superior ability to extract details, they struggle in geometrically comple…

Cited by 7SourcePDFScholar
2024

Box2Poly: Memory-Efficient Polygon Prediction of Arbitrarily Shaped and Rotated Text

AAAI 2024technical

Recently, Transformer-based text detection techniques have sought to predict polygons by encoding the coordinates of individual boundary vertices using distinct query features. However, this approach incurs a significant memory overhead and struggles to effectively capture the intricate relationship…

2024

DGInStyle: Domain-Generalizable Semantic Segmentation with Image Diffusion Models and Stylized Semantic Control

ECCV 2024poster

"Large, pretrained latent diffusion models (LDMs) have demonstrated an extraordinary ability to generate creative content, specialize to user data through few-shot fine-tuning, and condition their output on other modalities, such as semantic maps. However, are they usable as large-scale data generat…

2024

Dynamic LiDAR Re-simulation using Compositional Neural Fields

CVPR 2024highlight

We introduce DyNFL a novel neural field-based approach for high-fidelity re-simulation of LiDAR scans in dynamic driving scenes. DyNFL processes LiDAR measurements from dynamic environments accompanied by bounding boxes of moving objects to construct an editable neural field. This field comprising s…

2024

Living Scenes: Multi-object Relocalization and Reconstruction in Changing 3D Environments

CVPR 2024highlight

Research into dynamic 3D scene understanding has primarily focused on short-term change tracking from dense observations while little attention has been paid to long-term changes with sparse observations. We address this gap with MoRE a novel approach for multi-object relocalization and reconstructi…

2024

Point2CAD: Reverse Engineering CAD Models from 3D Point Clouds

CVPR 2024highlight

Computer-Aided Design (CAD) model reconstruction from point clouds is an important problem at the intersection of computer vision graphics and machine learning; it saves the designer significant time when iterating on in-the-wild objects. Recent advancements in this direction achieve relatively reli…

2024

Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation

CVPR 2024poster

Monocular depth estimation is a fundamental computer vision task. Recovering 3D depth from a single image is geometrically ill-posed and requires scene understanding so it is not surprising that the rise of deep learning has led to a breakthrough. The impressive progress of monocular depth estimator…

2024

StegoGAN: Leveraging Steganography for Non-Bijective Image-to-Image Translation

CVPR 2024poster

Most image-to-image translation models postulate that a unique correspondence exists between the semantic classes of the source and target domains. However this assumption does not always hold in real-world scenarios due to divergent distributions different class sets and asymmetrical information re…

2024

TetraDiffusion: Tetrahedral Diffusion Models for 3D Shape Generation

ECCV 2024poster

"Probabilistic denoising diffusion models (DDMs) have set a new standard for 2D image generation. Extending DDMs for 3D content creation is an active field of research. Here, we propose TetraDiffusion, a diffusion model that operates on a tetrahedral partitioning of 3D space to enable efficient, hig…

2023

BITE: Beyond Priors for Improved Three-D Dog Pose Estimation

CVPR 2023poster

We address the problem of inferring the 3D shape and pose of dogs from images. Given the lack of 3D training data, this problem is challenging, and the best methods lag behind those designed to estimate human shape and pose. To make progress, we attack the problem from multiple sides at once. First,…

Cited by 29SourcePDFScholar
2023

BiasBed - Rigorous Texture Bias Evaluation

CVPR 2023poster

The well-documented presence of texture bias in modern convolutional neural networks has led to a plethora of algorithms that promote an emphasis on shape cues, often to support generalization to new domains. Yet, common datasets, benchmarks and general model selection strategies are missing, and th…

2023

Connecting the Dots: Floorplan Reconstruction Using Two-Level Queries

CVPR 2023poster

We address 2D floorplan reconstruction from 3D scans. Existing approaches typically employ heuristically designed multi-stage pipelines. Instead, we formulate floorplan reconstruction as a single-stage structured prediction task: find a variable-size set of polygons, which in turn are variable-lengt…

2023

Guided Depth Super-Resolution by Deep Anisotropic Diffusion

CVPR 2023poster

Performing super-resolution of a depth image using the guidance from an RGB image is a problem that concerns several fields, such as robotics, medical imaging, and remote sensing. While deep learning methods have achieved good results in this problem, recent work highlighted the value of combining m…

2023

Neural LiDAR Fields for Novel View Synthesis

ICCV 2023poster

We present Neural Fields for LiDAR (NFL), a method to optimise a neural field scene representation from LiDAR measurements, with the goal of synthesizing realistic LiDAR scans from novel viewpoints. NFL combines the rendering power of neural fields with a detailed, physically motivated model of the…

Cited by 61PDFScholar
2022

BARC: Learning To Regress 3D Dog Shape From Images by Exploiting Breed Information

CVPR 2022poster

Our goal is to recover the 3D shape and pose of dogs from a single image. This is a challenging task because dogs exhibit a wide range of shapes and appearances, and are highly articulated. Recent work has proposed to directly regress the SMAL animal model, with additional limb scale parameters, fro…

Cited by 52PDFScholar
2022

Dynamic 3D Scene Analysis by Point Cloud Accumulation

ECCV 2022poster

"Multi-beam LiDAR sensors, as used on autonomous vehicles and mobile robots, acquire sequences of 3D range scans (""frames""). Each frame covers the scene sparsely, due to limited angular scanning resolution and occlusion. The sparsity restricts the performance of downstream processes like semantic…

2022

FiLM-Ensemble: Probabilistic Deep Learning via Feature-wise Linear Modulation

NeurIPS 2022accept

The ability to estimate epistemic uncertainty is often crucial when deploying machine learning in the real world, but modern methods often produce overconfident, uncalibrated uncertainty predictions. A common approach to quantify epistemic uncertainty, usable across a wide class of prediction models…

2022

Learning Graph Regularisation for Guided Super-Resolution

CVPR 2022poster

We introduce a novel formulation for guided super-resolution. Its core is a differentiable optimisation layer that operates on a learned affinity graph. The learned graph potentials make it possible to leverage rich contextual information from the guide image, while the explicit graph optimisation w…

Cited by 47PDFcodeScholar
2021

Cherry-Picking Gradients: Learning Low-Rank Embeddings of Visual Data via Differentiable Cross-Approximation

ICCV 2021poster

We propose an end-to-end trainable framework that processes large-scale visual data tensors by looking at a fraction of their entries only. Our method combines a neural network encoder with a tensor train decomposition to learn a low-rank latent encoding, coupled with cross-approximation (CA) to lea…

Cited by 6PDFcodeScholar
2021

In the Light of Feature Distributions: Moment Matching for Neural Style Transfer

CVPR 2021poster

Style transfer aims to render the content of a given image in the graphical/artistic style of another image. The fundamental concept underlying Neural Style Transfer (NST) is to interpret style as a distribution in the feature space of a Convolutional Neural Network, such that a desired style can be…

Cited by 60PDFcodeScholar
2021

PC2WF: 3D Wireframe Reconstruction from Raw Point Clouds

ICLR 2021poster

We introduce PC2WF, the first end-to-end trainable deep network architecture to convert a 3D point cloud into a wireframe model. The network takes as input an unordered set of 3D points sampled from the surface of some object, and outputs a wireframe of that object, i.e., a sparse set of corner poin…

Cited by 49SourcePDFScholar
2021

Predator: Registration of 3D Point Clouds With Low Overlap

CVPR 2021poster

We introduce PREDATOR, a model for pairwise pointcloud registration with deep attention to the overlap region. Different from previous work, our model is specifically designed to handle (also) point-cloud pairs with low overlap. Its key novelty is an overlap-attention block for early information exc…

Cited by 646PDFcodeScholar
2020

From Two Rolling Shutters to One Global Shutter

CVPR 2020oral

Most consumer cameras are equipped with electronic rolling shutter, leading to image distortions when the camera moves during image capture. We explore a surprisingly simple camera configuration that makes it possible to undo the rolling shutter distortion: two cameras mounted to have different roll…

Cited by 41PDFScholar
2020

Minimal Rolling Shutter Absolute Pose with Unknown Focal Length and Radial Distortion

ECCV 2020poster

The internal geometry of most modern consumer cameras is not adequately described by the perspective projection. Almost all cameras exhibit some radial lens distortion and are equipped with electronic rolling shutter that induces distortions when the camera moves during the image capture. When focal…

2020

Reconstruction of 3D flight trajectories from ad-hoc camera networks

IROS 2020poster

We present a method to reconstruct the 3D trajectory of an airborne robotic system only from videos recorded with cameras that are unsynchronized, may feature rolling shutter distortion, and whose viewpoints are unknown. Our approach enables robust and accurate outside-in tracking of dynamically fly…

Cited by 19SourceScholar
2020

Reconstruction of 3D ight trajectories from ad-hoc camera networks

IROS 2020

We present a method to reconstruct the 3D trajectory of an airborne robotic system only from videos recorded with cameras that are unsynchronized, may feature rolling shutter distortion, and whose viewpoints are unknown. Our approach enables robust and accurate outside-in tracking of dynamically fly

Cited by 25SourceScholar
2019

Guided Super-Resolution As Pixel-to-Pixel Transformation

ICCV 2019poster

Guided super-resolution is a unifying framework for several computer vision tasks where the inputs are a low-resolution source image of some target quantity (e.g., perspective depth acquired with a time-of-flight camera) and a high-resolution guide image from a different domain (e.g., a grey-scale i…

Cited by 90PDFcodeScholar
2017

A Multi-View Stereo Benchmark With High-Resolution Images and Multi-Camera Videos

CVPR 2017poster

Motivated by the limitations of existing multi-view stereo benchmarks, we present a novel dataset for this task. Towards this goal, we recorded a variety of indoor and outdoor scenes using a high-precision laser scanner and captured both high-resolution DSLR imagery as well as synchronized low-resol…

Cited by 1009PDFScholar
2017

Semantically Informed Multiview Surface Refinement

ICCV 2017poster

We present a method to jointly refine the geometry and semantic segmentation of 3D surface meshes. Our method alternates between updating the shape and the semantic labels. In the geometry refinement step, the mesh is deformed with variational energy minimization, such that it simultaneously maximiz…

Cited by 37PDFScholar
2017

Volumetric Flow Estimation for Incompressible Fluids Using the Stationary Stokes Equations

ICCV 2017poster

In experimental fluid dynamics, the flow in a volume of fluid is observed by injecting high-contrast tracer particles and tracking them in multi-view video. Fluid dynamics researchers have developed variants of space-carving to reconstruct the 3D particle distribution at a given time-step, and then…

Cited by 13PDFScholar
2016

Cataloging Public Objects Using Aerial and Street-Level Images - Urban Trees

CVPR 2016accepted

Each corner of the inhabited world is imaged from multiple viewpoints with increasing frequency. Online map services like Google Maps or Here Maps provide direct access to huge amounts of densely sampled, georeferenced images from street view and aerial perspective. There is an opportunity to design…

Cited by 201SourcePDFScholar
2016

Just Look at the Image: Viewpoint-Specific Surface Normal Prediction for Improved Multi-View Reconstruction

CVPR 2016spotlight

We present a multi-view reconstruction method that combines conventional multi-view stereo (MVS) with appearance-based normal prediction, to obtain dense and accurate 3D surface models. Reliable surface normals reconstructed from multi-view correspondence serve as training data for a convolutional n…

Cited by 32PDFScholar
2016

Large-Scale Location Recognition and the Geometric Burstiness Problem

CVPR 2016spotlight

Visual location recognition is the task of determining the place depicted in a query image from a given database of geo-tagged images. Location recognition is often cast as an image retrieval problem and recent research has almost exclusively focused on improving the chance that a relevant database…

Cited by 191PDFcodeScholar
2016

Large-Scale Semantic 3D Reconstruction: An Adaptive Multi-Resolution Model for Multi-Class Volumetric Labeling

CVPR 2016oral

We propose an adaptive multi-resolution formulation of semantic 3D reconstruction. Given a set of images of a scene, semantic 3D reconstruction aims to densely reconstruct both the 3D shape of the scene and a segmentation into semantic object classes. Jointly reasoning about shape and class allows o…

Cited by 119PDFScholar
2015

Hyperpoints and Fine Vocabularies for Large-Scale Location Recognition

ICCV 2015poster

Structure-based localization is the task of finding the absolute pose of a given query image w.r.t. a pre-computed 3D model. While this is almost trivial at small scale, special care must be taken as the size of the 3D model grows, because straight-forward descriptor matching becomes ineffective due…

Cited by 200PDFScholar
2015

Massively Parallel Multiview Stereopsis by Surface Normal Diffusion

ICCV 2015poster

We present a new, massively parallel method for high-quality multiview matching. Our work builds on the Patchmatch idea: starting from randomly generated 3D planes in scene space, the best-fitting planes are iteratively propagated and refined to obtain a 3D depth and normal field per view, such that…

Cited by 685PDFcodeScholar