← Search

Torsten Sattler

75 accepted papers

2025

Gaussian Splatting Feature Fields for (Privacy-Preserving) Visual Localization

CVPR 2025poster

Visual localization is the task of estimating a camera pose in a known environment. In this paper, we utilize 3D Gaussian Splatting (3DGS)-based representations for accurate and privacy-preserving visual localization. We propose Gaussian Splatting Feature Fields (GSFFs), a scene representation for v…

Cited by 0SourcePDFScholar
2025

LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering

NeurIPS 2025spotlight

In this work, we present a novel level-of-detail (LOD) method for 3D Gaussian Splatting that enables real-time rendering of large-scale scenes on memory-constrained devices. Our approach introduces a hierarchical LOD representation that iteratively selects optimal subsets of Gaussians based on camer…

Cited by 0SourceScholar
2025

Practical Solutions to the Relative Pose of Three Calibrated Cameras

CVPR 2025poster

We study the challenging problem of estimating the relative pose of three calibrated cameras from four point correspondences. We propose novel efficient solutions to this problem that are based on the simple idea of using four correspondences to estimate an approximate geometry of the first two view…

2025

RePoseD: Efficient Relative Pose Estimation With Known Depth Information

ICCV 2025poster

Recent advances in monocular depth estimation methods (MDE) and their improved accuracy open new possibilities for their applications. In this paper, we investigate how monocular depth estimates can be used for relative pose estimation. In particular, we are interested in answering the question whet…

2024

Absolute Pose from One or Two Scaled and Oriented Features

CVPR 2024highlight

Keypoints used for image matching often include an estimate of the feature scale and orientation. While recent work has demonstrated the advantages of using feature scales and orientations for relative pose estimation relatively little work has considered their use for absolute pose estimation. We i…

2024

Mip-Splatting: Alias-free 3D Gaussian Splatting

CVPR 2024poster

Recently 3D Gaussian Splatting has demonstrated impressive novel view synthesis results reaching high fidelity and efficiency. However strong artifacts can be observed when changing the sampling rate e.g. by changing focal length or camera distance. We find that the source for this phenomenon can be…

Cited by 345SourcePDFScholar
2024

The Unreasonable Effectiveness of Pre-Trained Features for Camera Pose Refinement

CVPR 2024highlight

Pose refinement is an interesting and practically relevant research direction. Pose refinement can be used to (1) obtain a more accurate pose estimate from an initial prior (e.g. from retrieval) (2) as pre-processing i.e. to provide a better starting point to a more expensive pose estimator (3) as p…

Cited by 5SourcePDFScholar
2024

WildGaussians: 3D Gaussian Splatting In the Wild

NeurIPS 2024poster

While the field of 3D scene reconstruction is dominated by NeRFs due to their photorealistic quality, 3D Gaussian Splatting (3DGS) has recently emerged, offering similar quality with real-time rendering speeds. However, both methods primarily excel with well-controlled 3D scenes, while in-the-wild d…

2023

P1AC: Revisiting Absolute Pose From a Single Affine Correspondence

ICCV 2023oral

Affine correspondences have traditionally been used to improve feature matching over wide baselines. While recent work has successfully used affine correspondences to solve various relative camera pose estimation problems, less attention has been given to their use in absolute pose estimation. We in…

Cited by 10PDFcodeScholar
2023

Privacy-Preserving Representations Are Not Enough: Recovering Scene Content From Camera Poses

CVPR 2023poster

Visual localization is the task of estimating the camera pose from which a given image was taken and is central to several 3D computer vision applications. With the rapid growth in the popularity of AR/VR/MR devices and cloud-based applications, privacy issues are becoming a very important aspect of…

2023

SegLoc: Learning Segmentation-Based Representations for Privacy-Preserving Visual Localization

CVPR 2023poster

Inspired by properties of semantic segmentation, in this paper we investigate how to leverage robust image segmentation in the context of privacy-preserving visual localization. We propose a new localization framework, SegLoc, that leverages image segmentation to create robust, compact, and privacy-…

Cited by 17SourcePDFScholar
2023

Visual Localization Using Imperfect 3D Models From the Internet

CVPR 2023poster

Visual localization is a core component in many applications, including augmented reality (AR). Localization algorithms compute the camera pose of a query image w.r.t. a scene representation, which is typically built from images. This often requires capturing and storing large amounts of data, follo…

2022

Deep Visual Geo-Localization Benchmark

CVPR 2022oral

In this paper, we propose a new open-source benchmarking framework for Visual Geo-localization (VG) that allows to build, train, and test a wide range of commonly used architectures, with the flexibility to change individual components of a geo-localization pipeline. The purpose of this framework is…

Cited by 103PDFcodeScholar
2022

MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction

NeurIPS 2022accept

In recent years, neural implicit surface reconstruction methods have become popular for multi-view 3D reconstruction. In contrast to traditional multi-view stereo methods, these approaches tend to produce smoother and more complete reconstructions due to the inductive smoothness bias of neural netwo…

Cited by 505SourcePDFScholar
2022

Objects Can Move: 3D Change Detection by Geometric Transformation Consistency

ECCV 2022poster

"AR/VR applications and robots need to know when the scene has changed. An example is when objects are moved, added, or removed from the scene. We propose a 3D object discovery method that is based only on scene changes. Our method does not need to encode any assumptions about what is an object, but…

2022

ViewFormer: NeRF-Free Neural Rendering from Few Images Using Transformers

ECCV 2022poster

"Novel view synthesis is a long-standing problem. In this work, we consider a variant of the problem where we are given only a few context views sparsely covering a scene or an object. The goal is to predict novel viewpoints in the scene, which requires learning priors. The current state of the art…

2021

Back to the Feature: Learning Robust Camera Localization From Pixels To Pose

CVPR 2021poster

Camera pose estimation in known scenes is a 3D geometry task recently tackled by multiple learning algorithms. Many regress precise geometric quantities, like poses or 3D points, from an input image. This either fails to generalize to new viewpoints or ties the model parameters to a specific scene.…

Cited by 301PDFcodeScholar
2021

Calibrated and Partially Calibrated Semi-Generalized Homographies

ICCV 2021poster

In this paper, we propose the first minimal solutions for estimating the semi-generalized homography given a perspective and a generalized camera. The proposed solvers use five 2D-2D image point correspondences induced by a scene plane. One group of solvers assumes the perspective camera to be fully…

Cited by 13PDFcodeScholar
2021

CrowdDriven: A New Challenging Dataset for Outdoor Visual Localization

ICCV 2021poster

Visual localization is the problem of estimating the position and orientation from which a given image (or a sequence of images) is taken in a known scene. It is an important part of a wide range of computer vision and robotics applications, from self-driving cars to augmented/virtual reality system…

Cited by 21PDFScholar
2021

How Privacy-Preserving Are Line Clouds? Recovering Scene Details From 3D Lines

CVPR 2021poster

Visual localization is the problem of estimating the camera pose of a given image with respect to a known scene. Visual localization algorithms are a fundamental building block in advanced computer vision applications, including Mixed and Virtual Reality systems. Many algorithms used in practice rep…

Cited by 29PDFcodeScholar
2021

Human POSEitioning System (HPS): 3D Human Pose Estimation and Self-Localization in Large Scenes From Body-Mounted Sensors

CVPR 2021poster

We introduce (HPS) Human POSEitioning System, a method to recover the full 3D pose of a human registered with a 3D scan of the surrounding environment using wearable sensors. Using IMUs attached at the body limbs and a head mounted camera looking outwards, HPS fuses camera based self-localization wi…

Cited by 171PDFcodeScholar
2021

On the Limits of Pseudo Ground Truth in Visual Camera Re-Localisation

ICCV 2021poster

Benchmark datasets that measure camera pose accuracy have driven progress in visual re-localisation research. To obtain poses for thousands of images, it is common to use a reference algorithm to generate pseudo ground truth. Popular choices include Structure-from-Motion (SfM) and Simultaneous-Local…

Cited by 75PDFcodeScholar
2020

Beyond Controlled Environments: 3D Camera Re-Localization in Changing Indoor Scenes

ECCV 2020poster

Long-term camera re-localization is an important task with numerous computer vision and robotics applications. Whilst various outdoor benchmarks exist that target lighting, weather and seasonal changes, far less attention has been paid to appearance changes that occur indoors. This has led to a mism…

2020

Handcrafted Outlier Detection Revisited

ECCV 2020poster

Local feature matching is a critical part of many computer vision pipelines, including among others Structure-from-Motion, SLAM, and Visual Localization. However, due to limitations in the descriptors, raw matches are often contaminated by a majority of outliers. As a result, outlier detection is a…

2020

Infrastructure-based Multi-Camera Calibration using Radial Projections

ECCV 2020poster

Multi-camera systems are an important sensor platform for intelligent systems such as self-driving cars. Pattern-based calibration techniques can be used to calibrate the intrinsics of the cameras individually. However, extrinsic calibration of systems with little to no visual overlap between the ca…

2020

Making Affine Correspondences Work in Camera Geometry Computation

ECCV 2020poster

Local features such as SIFT and its affine and learned variants provide region-to-region rather than point-to-point correspondences. It has recently been exploited to create new minimal solvers for classical problems such as homography, essential and fundamental matrix estimation. The main argument…

2020

Single-Image Depth Prediction Makes Feature Matching Easier

ECCV 2020poster

Good local features improve the robustness of many 3D re-localization and multi-view reconstruction pipelines. The problem is that viewing angle and distance severely impact the recognizability of a local feature. Attempts to improve appearance invariance by choosing better local feature points or b…

2020

To Learn or Not to Learn: Visual Localization from Essential Matrices

ICRA 2020poster

Visual localization is the problem of estimating a camera within a scene and a key technology for autonomous robots. State-of-the-art approaches for accurate visual localization use scene-specific representations, resulting in the overhead of constructing these models when applying the techniques to…

Cited by 128SourceScholar
2020

Why Having 10,000 Parameters in Your Camera Model Is Better Than Twelve

CVPR 2020oral

Camera calibration is an essential first step in setting up 3D Computer Vision systems. Commonly used parametric camera models are limited to a few degrees of freedom and thus often do not optimally fit to complex real lens distortion. In contrast, generic camera models allow for very accurate calib…

Cited by 70PDFcodeScholar
2019

A Cross-Season Correspondence Dataset for Robust Semantic Segmentation

CVPR 2019poster

In this paper, we present a method to utilize 2D-2D point matches between images taken during different image conditions to train a convolutional neural network for semantic segmentation. Enforcing label consistency across the matches makes the final segmentation algorithm robust to seasonal changes…

Cited by 104PDFcodeScholar
2019

D2-Net: A Trainable CNN for Joint Description and Detection of Local Features

CVPR 2019poster

In this work we address the problem of finding reliable pixel-level correspondences under difficult imaging conditions. We propose an approach where a single convolutional neural network plays a dual role: It is simultaneously a dense feature descriptor and a feature detector. By postponing the dete…

Cited by 909PDFcodeScholar
2019

Efficient 2D-3D Matching for Multi-Camera Visual Localization

ICRA 2019poster

Visual localization, i.e., determining the position and orientation of a vehicle with respect to a map, is a key problem in autonomous driving. We present a multi-camera visual inertial localization algorithm for large scale environments. To efficiently and effectively match features against a pre-b…

Cited by 41SourceScholar
2019

Fine-Grained Segmentation Networks: Self-Supervised Segmentation for Improved Long-Term Visual Localization

ICCV 2019poster

Long-term visual localization is the problem of estimating the camera pose of a given query image in a scene whose appearance changes over time. It is an important problem in practice that is, for example, encountered in autonomous driving. In order to gain robustness to such changes, long-term loca…

Cited by 88PDFcodeScholar
2019

Incremental Visual-Inertial 3D Mesh Generation with Structural Regularities

ICRA 2019poster

Visual-Inertial Odometry (VIO) algorithms typically rely on a point cloud representation of the scene that does not model the topology of the environment. A 3D mesh instead offers a richer, yet lightweight, model. Nevertheless, building a 3D mesh out of the sparse and noisy 3D landmarks triangulated…

Cited by 64SourcecodeScholar
2019

Is This the Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization

ICCV 2019poster

Visual localization in large and complex indoor scenes, dominated by weakly textured rooms and repeating geometric patterns, is a challenging problem with high practical relevance for applications such as Augmented Reality and robotics. To handle the ambiguities arising in this scenario, a common st…

Cited by 59PDFScholar
2019

Night-to-Day Image Translation for Retrieval-based Localization

ICRA 2019poster

Visual localization is a key step in many robotics pipelines, allowing the robot to (approximately) determine its position and orientation in the world. An efficient and scalable approach to visual localization is to use image retrieval techniques. These approaches identify the image most similar to…

Cited by 271SourcecodeScholar
2019

Project AutoVision: Localization and 3D Scene Perception for an Autonomous Vehicle with a Multi-Camera System

ICRA 2019poster

Project AutoVision aims to develop localization and 3D scene perception capabilities for a self-driving vehicle. Such capabilities will enable autonomous navigation in urban and rural environments, in day and night, and with cameras as the only exteroceptive sensors. The sensor suite employs many ca…

Cited by 149SourceScholar
2019

Real-Time Dense Mapping for Self-Driving Vehicles using Fisheye Cameras

ICRA 2019poster

We present a real-time dense geometric mapping algorithm for large-scale environments. Unlike existing methods which use pinhole cameras, our implementation is based on fisheye cameras whose large field of view benefits various computer vision applications for self-driving vehicles such as visual-in…

Cited by 47SourceScholar
2019

Understanding the Limitations of CNN-Based Absolute Camera Pose Regression

CVPR 2019poster

Visual localization is the task of accurate camera pose estimation in a known scene. It is a key problem in computer vision and robotics, with applications including self-driving cars, Structure-from-Motion, SLAM, and Mixed Reality. Traditionally, the localization problem has been tackled using 3D g…

Cited by 469PDFcodeScholar
2018

Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions

CVPR 2018poster

Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing condition, including day-night changes, as well as weather and…

Cited by 780SourcePDFScholar
2018

InLoc: Indoor Visual Localization With Dense Matching and View Synthesis

CVPR 2018poster

We seek to predict the 6 degree-of-freedom (6DoF) pose of a query photograph with respect to a large indoor 3D map. The contributions of this work are three-fold. First, we develop a new large-scale visual localization method targeted for indoor environments. The method proceeds along three steps: (…

Cited by 582SourcePDFScholar
2018

Incremental Object Database: Building 3D Models from Multiple Partial Observations

IROS 2018poster

Collecting 3D object data sets involves a large amount of manual work and is time consuming. Getting complete models of objects either requires a 3D scanner that covers all the surfaces of an object or one needs to rotate it to completely observe it. We present a system that incrementally builds a d…

Cited by 48SourceScholar
2018

Semantic Match Consistency for Long-Term Visual Localization

ECCV 2018poster

Robust and accurate visual localization across large appearance variations due to changes in time of day, seasons, or changes of the environment is a challenging problem which is of importance to application areas such as navigation of autonomous robots. Traditional feature-based methods often strug…

Cited by 166SourcePDFScholar
2018

Towards Robust Visual Odometry with a Multi-Camera System

IROS 2018poster

We present a visual odometry (VO) algorithm for a multi-camera system and robust operation in challenging environments. Our algorithm consists of a pose tracker and a local mapper. The tracker estimates the current pose by minimizing photometric errors between the most recent keyframe and the curren…

Cited by 56SourceScholar
2018

VSO: Visual Semantic Odometry

ECCV 2018poster

Robust data association is a core problem of visual odometry, where image-to-image correspondences provide constraints for camera pose and map estimation. Current state-of-the-art direct and indirect methods use short-term tracking to obtain continuous frame-to-frame constraints, while long-term con…

Cited by 162SourcePDFScholar
2017

A Multi-View Stereo Benchmark With High-Resolution Images and Multi-Camera Videos

CVPR 2017poster

Motivated by the limitations of existing multi-view stereo benchmarks, we present a novel dataset for this task. Towards this goal, we recorded a variety of indoor and outdoor scenes using a high-precision laser scanner and captured both high-resolution DSLR imagery as well as synchronized low-resol…

Cited by 1009PDFScholar
2017

Are Large-Scale 3D Models Really Necessary for Accurate Visual Localization?

CVPR 2017poster

Accurate visual localization is a key technology for autonomous navigation. 3D structure-based methods employ 3D models of the scene to estimate the full 6DOF pose of a camera very accurately. However, constructing (and extending) large-scale 3D models is still a significant challenge. In contrast,…

Cited by 257PDFScholar
2017

Comparative Evaluation of Hand-Crafted and Learned Local Features

CVPR 2017poster

Matching local image descriptors is a key step in many computer vision applications. For more than a decade, hand-crafted descriptors such as SIFT have been used for this task. Recently, multiple new descriptors learned from data have been proposed and shown to improve on SIFT in terms of discrimina…

Cited by 381PDFScholar
2017

Direct visual odometry for a fisheye-stereo camera

IROS 2017poster

We present a direct visual odometry algorithm for a fisheye-stereo camera. Our algorithm performs simultaneous camera motion estimation and semi-dense reconstruction. The pipeline consists of two threads: a tracking thread and a mapping thread. In the tracking thread, we estimate the camera pose via…

Cited by 49SourceScholar
2017

Image-Based Localization Using LSTMs for Structured Feature Correlation

ICCV 2017poster

In this work we propose a new CNN+LSTM architecture for camera pose regression for indoor and outdoor scenes. CNNs allow us to learn suitable feature representations for localization that are robust against motion blur and illumination changes. We make use of LSTM units on the CNN output, which play…

Cited by 658PDFScholar
2017

Quad-Networks: Unsupervised Learning to Rank for Interest Point Detection

CVPR 2017poster

Several machine learning tasks require to represent the data using only a sparse set of interest points. An ideal detector is able to find the corresponding interest points even if the data undergo a transformation typical for a given domain. Since the task is of high practical interest in computer…

Cited by 230PDFScholar
2017

Semantically Informed Multiview Surface Refinement

ICCV 2017poster

We present a method to jointly refine the geometry and semantic segmentation of 3D surface meshes. Our method alternates between updating the shape and the semantic labels. In the geometry refinement step, the mesh is deformed with variational energy minimization, such that it simultaneously maximiz…

Cited by 37PDFScholar
2017

Toroidal Constraints for Two-Point Localization Under High Outlier Ratios

CVPR 2017poster

Localizing a query image against a 3D model at large scale is a hard problem, since 2D-3D matches become more and more ambiguous as the model size increases. This creates a need for pose estimation strategies that can handle very low inlier ratios. In this paper, we draw new insights on the geometri…

Cited by 42PDFScholar
2016

Large-Scale Location Recognition and the Geometric Burstiness Problem

CVPR 2016spotlight

Visual location recognition is the task of determining the place depicted in a query image from a given database of geo-tagged images. Location recognition is often cast as an image retrieval problem and recent research has almost exclusively focused on improving the chance that a relevant database…

Cited by 191PDFcodeScholar
2015

Get Out of My Lab: Large-scale, Real-Time Visual-Inertial Localization

RSS 2015poster

Accurately estimating a robot's pose relative to a global scene model and precisely tracking the pose in real-time is a fundamental problem for navigation and obstacle avoidance tasks. Due to the computational complexity of localization against a large map and the memory consumed by the model, state…

Cited by 306SourcePDFScholar
2015

Hyperpoints and Fine Vocabularies for Large-Scale Location Recognition

ICCV 2015poster

Structure-based localization is the task of finding the absolute pose of a given query image w.r.t. a pre-computed 3D model. While this is almost trivial at small scale, special care must be taken as the size of the 3D model grows, because straight-forward descriptor matching becomes ineffective due…

Cited by 200PDFScholar
2015

Non-Parametric Structure-Based Calibration of Radially Symmetric Cameras

ICCV 2015poster

We propose a novel two-step method for estimating the intrinsic and extrinsic calibration of any radially symmetric camera, including non-central systems. The first step consists of estimating the camera pose, given a Structure from Motion (SfM) model, up to the translation along the optical axis. A…

Cited by 22PDFScholar
2015

Obstacle detection for self-driving cars using only monocular cameras and wheel odometry

IROS 2015poster

Mapping the environment is crucial to enable path planning and obstacle avoidance for self-driving vehicles and other robots. In this paper, we concentrate on ground-based vehicles and present an approach which extracts static obstacles from depth maps computed out of multiple consecutive images. In…

Cited by 116SourceScholar
2015

Optimizing the Viewing Graph for Structure-From-Motion

ICCV 2015poster

The viewing graph represents a set of views that are related by pairwise relative geometries. In the context of Structure-from-Motion (SfM), the viewing graph is the input to the incremental or global estimation pipeline. Much effort has been put towards developing robust algorithms to overcome pote…

Cited by 165PDFScholar