← Search

Vladlen Koltun

100 accepted papers

2026

Sharp Monocular View Synthesis in Less Than a Second

ICLR 2026poster

We present SHARP, an approach to photorealistic view synthesis from a single image. Given a single photograph, SHARP regresses the parameters of a 3D Gaussian representation of the depicted scene. This is done in less than a second on a standard GPU via a single feedforward pass through a neural net…

Cited by 0SourcecodeScholar
2025

CoMotion: Concurrent Multi-person 3D Motion

ICLR 2025poster

We introduce an approach for detecting and tracking detailed 3D poses of multiple people from a single monocular camera stream. Our system maintains temporally coherent predictions in crowded scenes filled with difficult poses and occlusions. Our model performs both strong per-frame detection and a…

2025

Cut Your Losses in Large-Vocabulary Language Models

ICLR 2025oral

As language models grow ever larger, so do their vocabularies. This has shifted the memory footprint of LLMs during training disproportionately to one single layer: the cross-entropy in the loss computation. Cross-entropy builds up a logit matrix with entries for each pair of input tokens and vocabu…

2025

Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

ICLR 2025poster

We present a foundation model for zero-shot metric monocular depth estimation. Our model, Depth Pro, synthesizes high-resolution depth maps with unparalleled sharpness and high-frequency details. The predictions are metric, with absolute scale, without relying on the availability of metadata such as…

2025

Does Spatial Cognition Emerge in Frontier Models?

ICLR 2025poster

Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear when an organism traverses physical environments, smaller-s…

2025

Robust Autonomy Emerges from Self-Play

ICML 2025poster

Self-play has powered breakthroughs in two-player and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic driving emerges entirely from self-play in simulation at unprecedented scale -- 1.6 billion km of driving…

Cited by 4SourcePDFScholar
2024

OpenBot-Fleet: A System for Collective Learning with Real Robots

ICRA 2024poster

We introduce OpenBot-Fleet, a comprehensive open-source cloud robotics system for navigation. OpenBot-Fleet uses smartphones for sensing, local compute and communication, Google Firebase for secure cloud storage and off-board compute, and a robust yet low-cost wheeled robot to act in real-world envi…

Cited by 0SourceScholar
2022

Guaranteed Conservation of Momentum for Learning Particle-based Fluid Dynamics

NeurIPS 2022accept

We present a novel method for guaranteeing linear momentum in learned physics simulations. Unlike existing methods, we enforce conservation of momentum with a hard constraint, which we realize via antisymmetrical continuous convolutional layers. We combine these strict constraints with a hierarchica…

2022

Language-driven Semantic Segmentation

ICLR 2022poster

We present LSeg, a novel model for language-driven semantic image segmentation. LSeg uses a text encoder to compute embeddings of descriptive input labels (e.g., ``grass'' or ``building'') together with a transformer-based image encoder that computes dense per-pixel embeddings of the input image. Th…

2022

Shape From Polarization for Complex Scenes in the Wild

CVPR 2022poster

We present a new data-driven approach with physics-based priors to scene-level normal estimation from a single polarization image. Existing shape from polarization (SfP) works mainly focus on estimating the normal of a single object rather than complex scenes in the wild. A key barrier to high-quali…

Cited by 69PDFcodeScholar
2021

Deep Drone Acrobatics (Extended Abstract)

IJCAI 2021poster

Acrobatic flight with quadrotors is extremely challenging. Maneuvers such as the loop, matty flip, or barrel roll require high thrust and extreme angular accelerations that push the platform to its limits. Human drone pilots require years of practice to safely master such maneuvers. Yet, a tiny mis…

Cited by 0SourcePDFScholar
2021

Differentiable Simulation of Soft Multi-body Systems

NeurIPS 2021poster

We present a method for differentiable simulation of soft articulated bodies. Our work enables the integration of differentiable physical dynamics into gradient-based pipelines. We develop a top-down matrix assembly algorithm within Projective Dynamics and derive a generalized dry friction model for…

2021

Efficient Differentiable Simulation of Articulated Bodies

ICML 2021spotlight

We present a method for efficient differentiable simulation of articulated bodies. This enables integration of articulated body dynamics into deep learning frameworks, and gradient-based optimization of neural networks that operate on articulated bodies. We derive the gradients of the contact solver…

2021

Habitat 2.0: Training Home Assistants to Rearrange their Habitat

NeurIPS 2021spotlight

We introduce Habitat 2.0 (H2.0), a simulation platform for training virtual robots in interactive 3D environments and complex physics-enabled scenarios. We make comprehensive contributions to all levels of the embodied AI stack – data, simulation, and benchmark tasks. Specifically, we present: (i) R…

2021

Large Batch Simulation for Deep Reinforcement Learning

ICLR 2021poster

We accelerate deep reinforcement learning-based training in visually complex 3D environments by two orders of magnitude over prior work, realizing end-to-end training speeds of over 19,000 frames of experience per second on a single GPU and up to 72,000 frames per second on a single eight-GPU machin…

2021

Megaverse: Simulating Embodied Agents at One Million Experiences per Second

ICML 2021spotlight

We present Megaverse, a new 3D simulation platform for reinforcement learning and embodied AI research. The efficient design of our engine enables physics-based simulation with high-dimensional egocentric observations at more than 1,000,000 actions per second on a single 8-GPU node. Megaverse is up…

2021

Online Continual Learning With Natural Distribution Shifts: An Empirical Study With Visual Data

ICCV 2021poster

Continual learning is the problem of learning and retaining knowledge through time over multiple tasks and environments. Research has primarily focused on the incremental classification setting, where new tasks/classes are added at discrete time intervals. Such an "offline" setting does not evaluate…

Cited by 106PDFcodeScholar
2021

Stable View Synthesis

CVPR 2021poster

We present Stable View Synthesis (SVS). Given a set of source images depicting a scene from freely distributed viewpoints, SVS synthesizes new views of the scene. The method operates on a geometric scaffold computed via structure-from-motion and multi-view stereo. Each point on this 3D scaffold is a…

Cited by 223PDFcodeScholar
2021

Training Graph Neural Networks with 1000 Layers

ICML 2021spotlight

Deep graph neural networks (GNNs) have achieved excellent results on various tasks on increasingly large graph datasets with millions of nodes and edges. However, memory complexity has become a major obstacle when training deep GNNs for practical applications due to the immense number of nodes, edge…

2020

Deep Drone Acrobatics

RSS 2020poster

Performing acrobatic maneuvers with quadrotors is extremely challenging. Acrobatic flight requires high thrust and extreme angular accelerations that push the platform to its physical limits. Professional drone pilots often measure their level of mastery by flying such maneuvers in competitions. In…

2020

Dynamic Low-light Imaging with Quanta Image Sensors

ECCV 2020poster

Imaging in low light is difficult because the number of photons arriving at the sensor is low. Imaging dynamic scenes in low-light environments is even more difficult because as the scene moves, pixels in adjacent frames need to be aligned before they can be denoised. Conventional CMOS image sensors…

Cited by 54SourcePDFScholar
2020

High-Dimensional Convolutional Networks for Geometric Pattern Recognition

CVPR 2020oral

High-dimensional geometric patterns appear in many computer vision problems. In this work, we present high-dimensional convolutional networks for geometric pattern recognition problems that arise in 2D and 3D registration problems. We first propose high-dimensional convolutional networks from 4 to 3…

Cited by 47PDFcodeScholar
2020

Lagrangian Fluid Simulation with Continuous Convolutions

ICLR 2020poster

We present an approach to Lagrangian fluid simulation with a new type of convolutional network. Our networks process sets of moving particles, which describe fluids in space and time. Unlike previous approaches, we do not build an explicit graph structure to connect the particles but use spatial con…

Cited by 229SourceScholar
2020

MSeg: A Composite Dataset for Multi-Domain Semantic Segmentation

CVPR 2020poster

We present MSeg, a composite dataset that unifies se- mantic segmentation datasets from different domains. A naive merge of the constituent datasets yields poor performance due to inconsistent taxonomies and annotation practices. We reconcile the taxonomies and bring the pixel-level annotations into…

Cited by 237PDFcodeScholar
2020

On Joint Estimation of Pose, Geometry and svBRDF From a Handheld Scanner

CVPR 2020poster

We propose a novel formulation for joint recovery of camera pose, object geometry and spatially-varying BRDF. The input to our approach is a sequence of RGB-D images captured by a mobile, hand-held scanner that actively illuminates the scene with point light sources. Compared to previous works that…

Cited by 55PDFScholar
2020

Sample Factory: Egocentric 3D Control from Pixels at 100000 FPS with Asynchronous Reinforcement Learning

ICML 2020poster

Increasing the scale of reinforcement learning experiments has allowed researchers to achieve unprecedented results in both training sophisticated agents for video games, and in sim-to-real transfer for robotics. Typically such experiments rely on large distributed systems and require expensive hard…

2019

Beauty and the Beast: Optimal Methods Meet Learning for Drone Racing

ICRA 2019poster

Autonomous micro aerial vehicles still struggle with fast and agile maneuvers, dynamic environments, imperfect sensing, and state estimation drift. Autonomous drone racing brings these challenges to the fore. Human pilots can fly a previously unseen track after a handful of practice runs. In contras…

Cited by 174SourceScholar
2019

Connecting the Dots: Learning Representations for Active Monocular Depth Estimation

CVPR 2019poster

We propose a technique for depth estimation with a monocular structured-light camera, i.e., a calibrated stereo set-up with one camera and one laser projector. Instead of formulating the depth estimation via a correspondence search problem, we show that a simple convolutional architecture is suffici…

Cited by 43PDFScholar
2019

Events-To-Video: Bringing Modern Computer Vision to Event Cameras

CVPR 2019poster

Event cameras are novel sensors that report brightness changes in the form of asynchronous "events" instead of intensity frames. They have significant advantages over conventional cameras: high temporal resolution, high dynamic range, and no motion blur. Since the output of event cameras is fundamen…

Cited by 460PDFScholar
2019

Habitat: A Platform for Embodied AI Research

ICCV 2019oral

We present Habitat, a platform for research in embodied artificial intelligence (AI). Habitat enables training embodied agents (virtual robots) in highly efficient photorealistic 3D simulation. Specifically, Habitat consists of: (i) Habitat-Sim: a flexible, high-performance 3D simulator with configu…

Cited by 2011PDFcodeScholar
2019

What Do Single-View 3D Reconstruction Networks Learn?

CVPR 2019poster

Convolutional networks for single-view object reconstruction have shown impressive performance and have become a popular subject of research. All existing techniques are united by the idea of having an encoder-decoder network that performs non-trivial reasoning about the 3D structure of the output s…

Cited by 521PDFScholar
2018

Combinatorial Optimization with Graph Convolutional Networks and Guided Tree Search

NeurIPS 2018poster

We present a learning-based approach to computing solutions for certain NP-hard problems. Our approach combines deep learning techniques with useful algorithmic elements from classic heuristics. The central component is a graph convolutional network that is trained to estimate the likelihood, for ea…

Cited by 637SourcePDFScholar
2018

Deep Drone Racing: Learning Agile Flight in Dynamic Environments

CoRL 2018

Autonomous agile flight brings up fundamental challenges in robotics, such as coping with unreliable state estimation, reacting optimally to dynamically changing environments, and coupling perception and action in real time under severe resource constraints. In this paper, we consider these challeng

Cited by 0SourcePDFScholar
2018

Driving Policy Transfer via Modularity and Abstraction

CoRL 2018

End-to-end approaches to autonomous driving have high sample complexity and are difficult to scale to realistic urban driving. Simulation can help end-to-end driving systems by providing a cheap, safe, and diverse training environment. Yet training driving policies in simulation brings up the proble

Cited by 0SourcePDFScholar
2018

End-to-End Driving Via Conditional Imitation Learning

ICRA 2018poster

Deep networks trained on demonstrations of human driving have learned to follow roads and avoid obstacles. However, driving policies trained via imitation learning cannot be controlled at test time. A vehicle trained end-to-end to imitate an expert cannot be guided to take a specific turn at an upco…

Cited by 1419SourcecodeScholar
2018

Motion Perception in Reinforcement Learning with Dynamic Objects

CoRL 2018

In dynamic environments, learned controllers are supposed to take motion into account when selecting the action to be taken. However, in existing reinforcement learning works motion is rarely treated explicitly; it is rather assumed that the controller learns the necessary motion representation from

Cited by 0SourcePDFScholar
2018

On Offline Evaluation of Vision-based Driving Models

ECCV 2018poster

Autonomous driving models should ideally be evaluated by deploying them on a fleet of physical vehicles in the real world. Unfortunately, this approach is not practical for the vast majority of researchers. An attractive alternative is to evaluate models offline, on a pre-collected validation datase…

2018

TD or not TD: Analyzing the Role of Temporal Differencing in Deep Reinforcement Learning

ICLR 2018poster

Our understanding of reinforcement learning (RL) has been shaped by theoretical and empirical results that were obtained decades ago using tabular representations and linear function approximators. These results suggest that RL methods that use temporal differencing (TD) are superior to direct Monte…

2018

Tangent Convolutions for Dense Prediction in 3D

CVPR 2018poster

We present an approach to semantic scene analysis using deep convolutional networks. Our approach is based on tangent convolutions - a new construction for convolutional networks on 3D data. In contrast to volumetric approaches, our method operates directly on surface geometry. Crucially, the constr…