← Search

Christian Rupprecht

69 accepted papers

2026

LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis

CVPR 2026

Recent work has shown that neural networks can perform 3D tasks such as Novel View Synthesis (NVS) without explicit 3D reconstruction. Even so, we argue that strong 3D inductive biases are still helpful in the design of such networks. We show this point by introducing LagerNVS, an encoder-decoder ne

Cited by 0SourcecodeScholar
2026

LitePT: Lighter Yet Stronger Point Transformer

CVPR 2026

Modern neural architectures for 3D point cloud processing contain both convolutional layers and attention blocks, but the best way to assemble them remains unclear. We analyse the role of different computational blocks in 3D point cloud networks and find an intuitive behaviour: convolution is adequa

Cited by 0SourcecodeScholar
2026

MaskInversion: Localized Embeddings via Optimization of Explainability Maps

ICLR 2026poster

Vision-language foundation models such as CLIP have achieved tremendous results in global vision-language alignment, but still show some limitations in creating representations for specific image regions. To address this problem, we propose MaskInversion, a method that leverages the feature represe…

Cited by 0SourcecodeScholar
2026

Particulate: Feed-Forward 3D Object Articulation

CVPR 2026

We introduce Particulate, a feed-forward model that, given a 3D mesh of an object, infers its articulations, including its 3D parts, their kinematic structure, and the motion constraints. The model is based on a transformer network, the Part Articulation Transformer, which predicts all these paramet

Cited by 0SourcecodeScholar
2026

S3OD: Towards Generalizable Salient Object Detection with Synthetic Data

ICLR 2026poster

Salient object detection exemplifies data-bounded tasks where expensive pixel-precise annotations force separate model training for related subtasks like DIS and HR-SOD. We present a method that dramatically improves generalization through large-scale synthetic data generation and ambiguity-aware ar…

Cited by 0SourcecodeScholar
2026

Scene-Centric Unsupervised Video Panoptic Segmentation

CVPR 2026

Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision. Existing unsupervised scene understanding works mainly focuse

Cited by 0SourceScholar
2026

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

CVPR 2026

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D understanding or video generation, a significant gap remains in

Cited by 0SourceScholar
2026

VFMF: Dense Forecasting by Generating Foundation Model Features

ICML 2026poster

Forecasting by generating RGB videos is computationally expensive, often physically implausible, and not directly actionable, since it requires translation into decision-making signals. Direct modality forecasting (e.g., predicting future segmentation) produces directly actionable outputs but fails …

Cited by 0SourceScholar
2026

What Happens Next? Anticipating Future Motion by Generating Point Trajectories

ICLR 2026poster

We consider the problem of forecasting motion from a single image, i.e., predicting how objects in the world are likely to move, without the ability to observe other parameters such as the object velocities or the forces applied to them. We formulate this task as conditional generation of dense traj…

Cited by 0SourcecodeScholar
2025

AnimalClue: Recognizing Animals by their Traces

ICCV 2025poster

Wildlife observation plays an important role in biodiversity conservation, necessitating robust methodologies for monitoring wildlife populations and interspecies interactions. Recent advances in computer vision have significantly contributed to automating fundamental wildlife observation tasks, suc…

2025

AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos

CVPR 2025poster

Estimating camera motion and intrinsics from casual videos is a core challenge in computer vision. Traditional bundle-adjustment based methods, such as SfM and SLAM, struggle to perform reliably on arbitrary data. Although specialized SfM approaches have been developed for handling dynamic scenes, t…

2025

CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance Shifts

ICCV 2025poster

An important challenge when using computer vision models in the real world is to evaluate their performance in potential out-of-distribution (OOD) scenarios. While simple synthetic corruptions are commonly applied to test OOD robustness, they often fail to capture nuisance shifts that occur in the r…

Cited by 0SourcePDFScholar
2025

CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

ICCV 2025poster

We introduce CoTracker3, a new state-of-the-art point tracker. With CoTracker3, we revisit the design of recent trackers, removing components and reducing the number of parameters while also improving performance. We also explore the interplay of synthetic and real data. Recent trackers are trained…

Cited by 0SourcePDFScholar
2025

DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness

ICCV 2025poster

Most 3D object generators prioritize aesthetic quality, often neglecting the physical constraints necessary for practical applications. One such constraint is that a 3D object should be self-supporting, i.e., remain balanced under gravity. Previous approaches to generating stable 3D objects relied o…

2025

Do It Yourself: Learning Semantic Correspondence from Pseudo-Labels

ICCV 2025poster

Finding correspondences between semantically similar points across images and object instances is one of the everlasting challenges in computer vision. While large pre-trained vision models have recently been demonstrated as effective priors for semantic matching, they still suffer from ambiguities…

Cited by 0SourcePDFScholar
2025

FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views

CVPR 2025poster

We present FLARE, a feed-forward model designed to infer high-quality camera poses and 3D geometry from uncalibrated sparse-view images (i.e., as few as 2-8 inputs), which is a challenging yet practical setting in real-world applications. Our solution features a cascaded learning paradigm with camer…

Cited by 0SourcePDFScholar
2025

Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion

ICCV 2025poster

Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC in an unsupervised setting. Our novel method, SceneDINO, adapts techniques from…

2025

Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics

ICCV 2025poster

We introduce Puppet-Master, an interactive video generator that captures the internal, part-level motion of objects, serving as a proxy for modeling object dynamics universally. Given an image of an object and a set of "drags" specifying the trajectory of a few points on the object, the model synthe…

Cited by 0SourcePDFScholar
2025

Scene-Centric Unsupervised Panoptic Segmentation

CVPR 2025highlight

Unsupervised panoptic segmentation aims to partition an image into semantically meaningful regions and distinct object instances without training on manually annotated data. In contrast to prior work on unsupervised panoptic scene understanding, we eliminate the need for object-centric training data…

2025

Stable Virtual Camera: Generative View Synthesis with Diffusion Models

ICCV 2025poster

We present \underline \text S tabl\underline \text e \underline \text V irtual C\underline \text a mera (Seva), a generalist diffusion model that creates novel views of a scene, given any number of input views and target cameras.Existing works struggle to generate either large viewpoint changes…

Cited by 0SourcePDFScholar
2025

SynCity: Training-Free Generation of 3D Worlds

ICCV 2025poster

We propose SynCity, a method for generating explorable 3D worlds from textual descriptions. Our approach leverages pre-trained textual, image, and 3D generators without requiring fine-tuning or inference-time optimization. While most 3D generators are object-centric and unable to create large-scale…

Cited by 0SourcePDFScholar
2025

Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks

CVPR 2025poster

We propose a new "Unbiased through Textual Description (UTD)" video benchmark based on unbiased subsets of existing video classification and retrieval datasets to enable a more robust assessment of video understanding capabilities. Namely, we tackle the problem that current video benchmarks may suff…

Cited by 0SourcePDFScholar
2025

VGGT: Visual Geometry Grounded Transformer

CVPR 2025award

We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typicall…

2024

Cache Me if You Can: Accelerating Diffusion Models through Block Caching

CVPR 2024poster

Diffusion models have recently revolutionized the field of image synthesis due to their ability to generate photorealistic images. However one of the major drawbacks of diffusion models is that the image generation process is costly. A large image-to-image network has to be applied many times to ite…

Cited by 51SourcePDFScholar
2024

CoTracker: It is Better to Track Together

ECCV 2024poster

"We introduce , a transformer-based model that tracks a large number of 2D points in long video sequences. Differently from most existing approaches that track points independently, tracks them jointly, accounting for their dependencies. We show that joint tracking significantly improves tracking ac…

2024

Diffusion Models for Open-Vocabulary Segmentation

ECCV 2024oral

"Open-vocabulary segmentation is the task of segmenting anything that can be named in an image. Recently, large-scale vision-language modelling has led to significant advances in open-vocabulary segmentation, but at the cost of gargantuan and increasing training and annotation efforts. Hence, we ask…

Cited by 8SourcePDFScholar
2024

DragAPart: Learning a Part-Level Motion Prior for Articulated Objects

ECCV 2024poster

"We introduce , a method that, given an image and a set of drags as input, generates a new image of the same object that responds to the action of the drags. Differently from prior works that focused on repositioning objects, predicts part-level interactions, such as opening and closing a drawer. We…

Cited by 14SourcePDFScholar
2024

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

ECCV 2024poster

"Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. However, in contrast to human-annotated captions, both speech and subtitles natu…

2024

IM-3D: Iterative Multiview Diffusion and Reconstruction for High-Quality 3D Generation

ICML 2024poster

Most text-to-3D generators build upon off-the-shelf text-to-image models trained on billions of images. They use variants of Score Distillation Sampling (SDS), which is slow, somewhat unstable, and prone to artifacts. A mitigation is to fine-tune the 2D generator to be multi-view aware, which can he…

Cited by 53SourcePDFScholar
2024

Learning Segmentation from Point Trajectories

NeurIPS 2024spotlight

We consider the problem of segmenting objects in videos based on their motion and no other forms of supervision. Prior work has often approached this problem by using the principle of common fate, namely the fact that the motion of points that belong to the same object is strongly correlated. Howeve…

2024

Learning the 3D Fauna of the Web

CVPR 2024poster

Learning 3D models of all animals in nature requires massively scaling up existing solutions. With this ultimate goal in mind we develop 3D-Fauna an approach that learns a pan-category deformable 3D animal model for more than 100 animal species jointly. One crucial bottleneck of modeling animals is…

Cited by 19SourcePDFScholar
2024

Rethinking Image Super Resolution from Training Data Perspectives

ECCV 2024poster

"In this work, we investigate the understudied effect of the training data used for image super-resolution (SR). Most commonly, novel SR methods are developed and benchmarked on common training datasets such as DIV2K and DF2K. However, we investigate and rethink the training data from the perspectiv…

2024

SHIC: Shape-Image Correspondences with no Keypoint Supervision

ECCV 2024poster

"Canonical surface mapping generalizes keypoint detection by assigning each pixel of an object to a corresponding point in a 3D template. Popularised by DensePose for the analysis of humans, authors have since attempted to apply the concept to more categories, but with limited success due to the hig…

Cited by 3SourcePDFScholar
2024

Scene-Conditional 3D Object Stylization and Composition

ECCV 2024poster

"Recently, 3D generative models have made impressive progress, enabling the generation of almost arbitrary 3D assets from text or image inputs. However, these approaches generate objects in isolation without any consideration for the scene where they will eventually be placed. In this paper, we prop…

Cited by 1SourcePDFScholar
2024

VGGSfM: Visual Geometry Grounded Deep Structure From Motion

CVPR 2024highlight

Structure-from-motion (SfM) is a long-standing problem in the computer vision community which aims to reconstruct the camera poses and 3D structure of a scene from a set of unconstrained 2D images. Classical frameworks solve this problem in an incremental manner by detecting and matching keypoints r…

2023

Behind the Scenes: Density Fields for Single View Reconstruction

CVPR 2023poster

Inferring a meaningful geometric scene representation from a single image is a fundamental problem in computer vision. Approaches based on traditional depth map prediction can only reason about areas that are visible in the image. Currently, neural radiance fields (NeRFs) can capture true 3D includi…

2023

Continual Detection Transformer for Incremental Object Detection

CVPR 2023poster

Incremental object detection (IOD) aims to train an object detector in phases, each with annotations for new object categories. As other incremental settings, IOD is subject to catastrophic forgetting, which is often addressed by techniques such as knowledge distillation (KD) and exemplar replay (ER…

Cited by 87SourcePDFScholar
2023

DynamicStereo: Consistent Dynamic Depth From Stereo Videos

CVPR 2023poster

We consider the problem of reconstructing a dynamic scene observed from a stereo camera. Most existing methods for depth from stereo treat different stereo frames independently, leading to temporally inconsistent depth predictions. Temporal consistency is especially important for immersive AR or VR…

2023

MagicPony: Learning Articulated 3D Animals in the Wild

CVPR 2023poster

We consider the problem of predicting the 3D shape, articulation, viewpoint, texture, and lighting of an articulated animal like a horse given a single test image as input. We present a new method, dubbed MagicPony, that learns this predictor purely from in-the-wild single-view images of the object…

2023

PC2: Projection-Conditioned Point Cloud Diffusion for Single-Image 3D Reconstruction

CVPR 2023highlight

Reconstructing the 3D shape of an object from a single RGB image is a long-standing problem in computer vision. In this paper, we propose a novel method for single-image 3D reconstruction which generates a sparse point cloud via a conditional denoising diffusion process. Our method takes as input a…

2023

PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment

ICCV 2023poster

Camera pose estimation is a long-standing computer vision problem that to date often relies on classical methods, such as handcrafted keypoint matching, RANSAC and bundle adjustment. In this paper, we propose to formulate the Structure from Motion (SfM) problem inside a probabilistic diffusion frame…

Cited by 83PDFcodeScholar
2023

RealFusion: 360deg Reconstruction of Any Object From a Single Image

CVPR 2023poster

We consider the problem of reconstructing a full 360deg photographic model of an object from a single image of it. We do so by fitting a neural radiance field to the image, but find this problem to be severely ill-posed. We thus take an off-the-self conditional image generator based on diffusion and…

Cited by 318SourcePDFScholar
2023

Temperature Schedules for self-supervised contrastive methods on long-tail data

ICLR 2023poster

Most approaches for self-supervised learning (SSL) are optimised on curated balanced datasets, e.g. ImageNet, despite the fact that natural data usually exhibits long-tail distributions. In this paper, we analyse the behaviour of one of the most popular variants of SSL, i.e. contrastive methods, on…

2023

Viewset Diffusion: (0-)Image-Conditioned 3D Generative Models from 2D Data

ICCV 2023poster

We present Viewset Diffusion, a diffusion-based generator that outputs 3D objects while only using multi-view 2D data for supervision. We note that there exists a one-to-one mapping between viewsets, i.e., collections of several 2D views of an object, and 3D models. Hence, we train a diffusion model…

Cited by 96PDFcodeScholar
2023

What does CLIP know about a red circle? Visual prompt engineering for VLMs

ICCV 2023oral

Large-scale Vision-Language Models, such as CLIP, learn powerful image-text representations that have found numerous applications, from zero-shot classification to text-to-image generation. Despite that, their capabilities for solving novel discriminative tasks via prompting fall behind those of lar…

Cited by 153PDFScholar
2022

Deep Spectral Methods: A Surprisingly Strong Baseline for Unsupervised Semantic Segmentation and Localization

CVPR 2022oral

Unsupervised localization and segmentation are long-standing computer vision challenges that involve decomposing an image into semantically-meaningful segments without any labeled data. These tasks are particularly interesting in an unsupervised setting due to the difficulty and cost of obtaining de…

Cited by 188PDFcodeScholar
2022

Finding an Unsupervised Image Segmenter in each of your Deep Generative Models

ICLR 2022poster

Recent research has shown that numerous human-interpretable directions exist in the latent space of GANs. In this paper, we develop an automatic procedure for finding directions that lead to foreground-background image separation, and we use these directions to train an image segmentation model with…

Cited by 62SourcePDFScholar
2022

Unsupervised Multi-Object Segmentation by Predicting Probable Motion Patterns

NeurIPS 2022accept

We propose a new approach to learn to segment multiple image objects without manual supervision. The method can extract objects form still images, but uses videos for supervision. While prior works have considered motion for segmentation, a key insight is that, while motion can be used to identify o…

Cited by 17SourcePDFScholar
2022

VTC: Improving Video-Text Retrieval with User Comments

ECCV 2022poster

"Multi-modal retrieval is an important problem for many applications, such as recommendation and search. Current benchmarks and even datasets are often manually constructed and consist of mostly clean samples where all modalities are well-correlated with the content. Thus, current video-text retriev…

2021

ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation

NeurIPS 2021poster

There has been a recent surge in methods that aim to decompose and segment scenes into multiple objects in an unsupervised manner, i.e., unsupervised multi-object segmentation. Performing such a task is a long-standing goal of computer vision, offering to unlock object-level reasoning without requir…

Cited by 84SourcecodeScholar
2021

Neural Response Interpretation Through the Lens of Critical Pathways

CVPR 2021poster

Is critical input information encoded in specific sparse pathways within the neural network? In this work, we discuss the problem of identifying these critical pathways and subsequently leverage them for interpreting the network's response to an input. The pruning objective --- selecting the smalles…

Cited by 41PDFcodeScholar
2021

PASS: An ImageNet replacement for self-supervised pretraining without humans

NeurIPS 2021poster

Computer vision has long relied on ImageNet and other large datasets of images sampled from the Internet for pretraining models. However, these datasets have ethical and technical shortcomings, such as containing personal information taken without consent, unclear license usage, biases, and, in some…

Cited by 67SourcecodeScholar
2021

Unsupervised Learning of Probably Symmetric Deformable 3D Objects from Images in the Wild (Extended Abstract)

IJCAI 2021poster

We propose a method to learn 3D deformable object categories from raw single-view images, without external supervision. The method is based on an autoencoder that factors each input image into depth, albedo, viewpoint and illumination. In order to disentangle these components without supervision, we…

2021

Unsupervised Part Discovery from Contrastive Reconstruction

NeurIPS 2021poster

The goal of self-supervised visual representation learning is to learn strong, transferable image representations, with the majority of research focusing on object or scene level. On the other hand, representation learning at part level has received significantly less attention. In this paper, we pr…

2020

Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents

ICLR 2020poster

As deep reinforcement learning driven by visual perception becomes more widely used there is a growing need to better understand and probe the learned agents. Understanding the decision making process and its relationship to visual inputs can be very valuable to identify problems in learned behavior…

Cited by 47SourceScholar
2020

Labelling unlabelled videos from scratch with multi-modal self-supervision

NeurIPS 2020poster

A large part of the current success of deep learning lies in the effectiveness of data -- more precisely: of labeled data. Yet, labelling a dataset with human annotation continues to carry high costs, especially for videos. While in the image domain, recent methods have allowed to generate meaningfu…

2020

Semantic Image Manipulation Using Scene Graphs

CVPR 2020poster

Image manipulation can be considered a special case of image generation where the image to be produced is a modification of an existing image. Image generation and manipulation have been, for the most part, tasks that operate on raw pixels. However, the remarkable progress in learning rich image and…

Cited by 146PDFcodeScholar
2020

Unsupervised Learning of Probably Symmetric Deformable 3D Objects From Images in the Wild

CVPR 2020oral

We propose a method to learn 3D deformable object categories from raw single-view images, without external supervision. The method is based on an autoencoder that factors each input image into depth, albedo, viewpoint and illumination. In order to disentangle these components without supervision, we…

Cited by 367PDFcodeScholar
2019

Explaining the Ambiguity of Object Detection and 6D Pose From Visual Data

ICCV 2019poster

3D object detection and pose estimation from a single image are two inherently ambiguous problems. Oftentimes, objects appear similar from different viewpoints due to shape symmetries, occlusion and repetitive textures. This ambiguity in both detection and pose estimation means that an object instan…

Cited by 138PDFScholar
2019

Towards Unsupervised Image Captioning With Shared Multimodal Embeddings

ICCV 2019poster

Understanding images without explicit supervision has become an important problem in computer vision. In this paper, we address image captioning by generating language descriptions of scenes without learning from annotated pairs of images and their captions. The core component of our approach is a s…

Cited by 144PDFcodeScholar
2018

Analyzing and Exploiting NARX Recurrent Neural Networks for Long-Term Dependencies

ICLR 2018workshop

Recurrent neural networks (RNNs) have achieved state-of-the-art performance on many diverse tasks, from machine translation to surgical activity recognition, yet training RNNs to capture long-term dependencies remains difficult. To date, the vast majority of successful RNN architectures alleviate th…

Cited by 36SourceScholar
2018

Guide Me: Interacting With Deep Networks

CVPR 2018poster

Interaction and collaboration between humans and intelligent machines has become increasingly important as machine learning methods move into real-world applications that involve end users. While much prior work lies at the intersection of natural language and vision, such as image captioning or ima…

Cited by 39SourcePDFScholar
2017

Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses

ICCV 2017poster

Many prediction tasks contain uncertainty. In some cases, uncertainty is inherent in the task itself. In future prediction, for example, many distinct outcomes are equally valid. In other cases, uncertainty arises from the way data is labeled. For example, in object detection, many objects of intere…

Cited by 235PDFScholar
2016

Sensor substitution for video-based action recognition

IROS 2016poster

There are many applications where domain-specific sensing, such as accelerometers, kinematics, or force sensing, provide unique and important information for control or for analysis of motion. However, it is not always the case that these sensors can be deployed or accessed beyond laboratory environ…

Cited by 34SourceScholar