← Search

Marc Pollefeys

252 accepted papers

2026

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

RSS 2026poster

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limi…

Cited by 0SourceScholar
2026

FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

CVPR 2026

We present FunREC, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunREC operates directly on in-t

Cited by 0SourcecodeScholar
2026

From Human Videos to Robot Manipulation: A Survey on Action-Relevant Representation Transfer for Scalable Vision-Language-Action Learning

IJCAI 2026

Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision–Language–Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly to obtain and tightly coupled to specific embodiments. Human vide

Cited by 0Scholar
2026

From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs

CVPR 2026

While Multimodal Large Language Models (MLLMs) have achieved impressive performance on semantic tasks, their spatial intelligence--crucial for robust and grounded AI systems--remains underdeveloped. Existing benchmarks fall short of diagnosing this limitation: they either focus on overly simplified

Cited by 0SourcecodeScholar
2026

FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning

CVPR 2026

Recent work in 3D scene understanding is moving beyond purely spatial analysis toward functional scene understanding. However, existing methods often consider functional relationships between object pairs in isolation, failing to capture the scene-wide interdependence that humans use to resolve ambi

Cited by 0SourcecodeScholar
2026

Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation

CVPR 2026

We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in 38 environments. Each object is operated in four embodiments

Cited by 0SourceScholar
2026

Loop Closure From Two Views: Revisiting PGO for Scalable Trajectory Estimation Through Monocular Priors

RA-L 2026

(Visual) Simultaneous Localization and Mapping (SLAM) remains a fundamental challenge in enabling autonomous systems to navigate and understand large-scale environments. Traditional SLAM approaches struggle to balance efficiency and accuracy, particularly in large-scale settings where extensive comp

Cited by 2SourceScholar
2026

MCGS-SLAM: A Multi-Camera SLAM Framework Using Gaussian Splatting for High-Fidelity Mapping

ICRA 2026poster

Recent progress in dense SLAM has primarily targeted monocular setups, often at the expense of robustness and geometric coverage. We present MCGS-SLAM, the first purely RGB-based multi-camera SLAM system built on 3D Gaussian Splatting (3DGS). Unlike prior methods relying on sparse maps or inertial d…

2026

Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision

ICML 2026poster

Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain, before they can reliably ground instructions and act. Yet, most available grounding datasets provide sparse supervision, with *insufficient* and *low-…

Cited by 0SourceScholar
2026

OVI-MAP: Open-Vocabulary Instance-Semantic Mapping

CVPR 2026

Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, real-time processing, and flexible open-set reasoning. Existing methods often rely

Cited by 0SourcecodeScholar
2026

OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

RSS 2026poster

Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely on dense 3D reconstruction and hand-crafted goal metrics, which limits their generalization across tasks and environmen…

Cited by 0SourceScholar
2026

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

ICML 2026poster

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures or necessitate compressing geometry into latent spaces to leverage pre-trained latent diffusion models. In this work, we demonstrate that such architectural overhead is unnecessary. We introduce a mini…

Cited by 0SourceScholar
2026

REACT3D: Recovering Articulations for Interactive Physical 3D Scenes

RA-L 2026

Interactive 3D scenes are increasingly vital for embodied intelligence, yet existing datasets remain limited due to the labor-intensive process of annotating part segmentation, kinematic types, and motion trajectories. We present REACT3D, a scalable zero-shot framework that converts static 3D scenes

Cited by 2SourceScholar
2026

SG2Loc: Sequential Visual Localization on 3D Scene Graphs

ICML 2026poster

Visual localization in complex environments remains a critical challenge for robotics and AR applications. Sequential localization, where pose estimates are refined over time, is important for autonomous agents. However, traditional methods often require storing extensive image databases or point cl…

Cited by 0SourceScholar
2026

Sight Over Site: Perception-Aware Reinforcement Learning for Efficient Robotic Inspection

ICRA 2026poster

Autonomous inspection is a central problem in robotics, with applications ranging from industrial monitoring to search-and-rescue. Traditionally, inspection has often been reduced to navigation tasks, where the objective is to reach a predefined location while avoiding obstacles. However, this formu…

2026

SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling

ICLR 2026poster

Generative methods for 3D assets have recently achieved remarkable progress, yet providing intuitive and precise control over the object geometry remains a key challenge. Existing approaches predominantly rely on text or image prompts, which often fall short in geometric specificity: language can be…

Cited by 0SourceScholar
2026

UnLoc: Leveraging Depth Uncertainties for Floorplan Localization

ICLR 2026poster

We propose UnLoc, an efficient data-driven solution for sequential camera localization within floorplans. Floorplan data is readily available, long-term persistent, and robust to changes in visual appearance. We address key limitations of recent methods, such as the lack of uncertainty modeling in d…

Cited by 0SourcecodeScholar
2026

Unblur-SLAM: Dense Neural SLAM for Blurry Inputs

CVPR 2026

We propose Unblur-SLAM, an RGB SLAM pipeline for sharp 3D reconstruction from blurred image inputs. In contrast to previous work, our approach is able to handle different types of blur and demonstrates state-of-the-art performance in the presence of both motion blur and defocus blur. Moreover, we ad

Cited by 0SourcecodeScholar
2026

YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting

ICLR 2026poster

Fast and flexible 3D scene reconstruction from unstructured image collections remains a significant challenge. We present YoNoSplat, a feedforward model that reconstructs high-quality 3D Gaussian Splatting representations from an arbitrary number of images. Our model is highly versatile, operating e…

Cited by 0SourceScholar
2025

3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection

ICCV 2025poster

Monocular 3D object detection is valuable for various applications such as robotics and AR/VR. Existing methods are confined to closed-set settings, where the training and testing sets consist of the same scenes and/or object categories. However, real-world applications often introduce new environme…

2025

ARKit LabelMaker: A New Scale for Indoor 3D Scene Understanding

CVPR 2025poster

Neural network performance scales with both model size and data volume, as shown in both language and image processing. This requires scaling-friendly architectures and large datasets. While transformers have been adapted for 3D vision, a `GPT-moment' remains elusive due to limited training data. We…

2025

ActLoc: Learning to Localize on the Move via Active Viewpoint Selection

CoRL 2025poster

Reliable localization is critical for robot navigation, yet many existing systems assume that all viewpoints along a trajectory are equally informative. In practice, localization becomes unreliable when the robot observes unmapped, ambiguous, or uninformative regions. To address this, we present Act…

Cited by 0SourceScholar
2025

Benchmarking Egocentric Visual-Inertial SLAM at City Scale

ICCV 2025poster

Precise 6-DoF simultaneous localization and mapping (SLAM) from onboard sensors is critical for wearable devices capturing egocentric data, which exhibits specific challenges, such as a wider diversity of motions and viewpoints, prevalent dynamic visual content, or long sessions affected by time-var…

Cited by 0SourcePDFScholar
2025

CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization

ICCV 2025poster

In dynamic 3D environments, accurately updating scene representations over time is crucial for applications in robotics, mixed reality, and embodied AI. As scenes evolve, efficient methods to incorporate changes are needed to maintain up-to-date, high-quality reconstructions without the computationa…

Cited by 0SourcePDFScholar
2025

CroCoDL: Cross-device Collaborative Dataset for Localization

CVPR 2025poster

Accurate localization plays a pivotal role in the autonomy of systems operating in unfamiliar environments, particularly when interaction with humans is expected. High-accuracy visual localization systems encompass various components, such as image retrievers, feature extractors, matchers, reconstru…

2025

CrossOver: 3D Scene Cross-Modal Alignment

CVPR 2025highlight

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene understanding via flexible, scene-level modality alignment.…

2025

DepthSplat: Connecting Gaussian Splatting and Depth

CVPR 2025poster

Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present DepthSplat to connect Gaussian splatting and depth estimation and study their interactions. More specifically, we first contribute a robust multi-view depth model by leveraging pre-train…

2025

EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models

NeurIPS 2025poster

Understanding behavior requires datasets that capture humans while carrying out complex tasks. The kitchen is an excellent environment for assessing human motor and cognitive function, as many complex actions are naturally exhibited in kitchens from chopping to cleaning. Here, we introduce the EPFL-…

Cited by 0SourcecodeScholar
2025

EgoM2P: Egocentric Multimodal Multitask Pretraining

ICCV 2025accepted

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the camera wearer's actions, intentions, and surrounding environ…

Cited by 0SourcePDFScholar
2025

EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision

CVPR 2025highlight

Touch contact and pressure are essential for understanding how humans interact with objects and offer insights that benefit applications in mixed reality and robotics. Estimating these interactions from an egocentric camera perspective is challenging, largely due to the lack of comprehensive dataset…

Cited by 3SourcePDFScholar
2025

FlowR: Flowing from Sparse to Dense 3D Reconstructions

ICCV 2025poster

3D Gaussian splatting enables high-quality novel view synthesis (NVS) at real-time frame rates. However, its quality drops sharply as we depart from the training views. Thus, dense captures are needed to match the high-quality expectations of applications like Virtual Reality (VR). However, such den…

Cited by 0SourcePDFScholar
2025

FrontierNet: Learning Visual Cues to Explore

RA-L 2025

Exploration of unknown environments is crucial for autonomous robots; it allows them to actively reason and decide on what new data to acquire for different tasks, such as mapping, object discovery, and environmental assessment. Existing solutions, such as frontier-based exploration approaches, rely

Cited by 12SourcecodeScholar
2025

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

CVPR 2025poster

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth o…

2025

Lost & Found: Tracking Changes From Egocentric Observations in 3D Dynamic Scene Graphs

RA-L 2025

Recent approaches have successfully focused on the segmentation of static reconstructions, thereby equipping downstream applications with semantic 3D understanding. However, the world in which we live is dynamic, characterized by numerous interactions between the environment and humans or robotic ag

Cited by 6SourcecodeScholar
2025

MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion

CVPR 2025poster

While Structure-from-Motion (SfM) has seen much progress over the years, state-of-the-art systems are prone to failure when facing extreme viewpoint changes in low-overlap, low-parallax or high-symmetry scenarios. Because capturing images that avoid these pitfalls is challenging, this severely limit…

2025

Multi-View 3D Point Tracking

ICCV 2025poster

We introduce the first data-driven multi-view 3D point tracker, designed to track arbitrary points in dynamic scenes using multiple camera views. Unlike existing monocular trackers, which struggle with depth ambiguities and occlusion, or prior multi-camera methods that require over 20 cameras and te…

2025

No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images

ICLR 2025oral

We introduce NoPoSplat, a feed-forward model capable of reconstructing 3D scenes parameterized by 3D Gaussians from unposed sparse multi-view images. Our model, trained exclusively with photometric loss, achieves real-time 3D Gaussian reconstruction during inference. To eliminate the need for accura…

2025

Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations

NeurIPS 2025poster

Learning effective multi-modal 3D representations of objects is essential for numerous applications, such as augmented reality and robotics. Existing methods often rely on task-specific embeddings that are tailored either for semantic understanding or geometric reconstruction. As a result, these e…

Cited by 0SourceScholar
2025

Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces

CVPR 2025highlight

We introduce the task of predicting functional 3D scene graphs for real-world indoor environments from posed RGB-D images. Unlike traditional 3D scene graphs that focus on spatial relationships of objects, functional 3D scene graphs capture objects, interactive elements, and their functional relatio…

2025

Planar Affine Rectification from Local Change of Scale and Orientation

ICCV 2025poster

We propose a method for affine rectification of an image plane by leveraging changes in local scales and orientations under projective distortion. Specifically, we derive a novel linear constraint that directly relates pairs of points with orientations to the parameters of a projective transformatio…

Cited by 0SourcePDFScholar
2025

R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization

CVPR 2025poster

Learning-based visual localization methods that use scene coordinate regression (SCR) offer the advantage of smaller map sizes. However, on datasets with complex illumination changes or image-level ambiguities, it remains a less robust alternative to feature matching methods. This work aims to close…

2025

Relative Pose Estimation through Affine Corrections of Monocular Depth Priors

CVPR 2025highlight

Monocular depth estimation (MDE) models have undergone significant advancements over recent years. Many MDE models aim to predict affine-invariant relative depth from monocular images, while recent developments in large-scale training and vision foundation models enable reasonable estimation of metr…

2025

Scaling Image Geo-Localization to Continent Level

NeurIPS 2025poster

Determining the precise geographic location of an image at a global scale remains an unsolved challenge. Standard image retrieval techniques are inefficient due to the sheer volume of images (>100M) and fail when coverage is insufficient. Scalable solutions, however, involve a trade-off: global cla…

Cited by 0SourcecodeScholar
2025

Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model

NeurIPS 2025poster

Physical environments and circumstances are fundamentally dynamic, yet current 3D datasets and evaluation benchmarks tend to concentrate on either dynamic scenarios or dynamic situations in isolation, resulting in incomplete comprehension. To overcome these constraints, we introduce Situat3DChange,…

Cited by 0SourcecodeScholar
2025

Structure-from-Motion with a Non-Parametric Camera Model

CVPR 2025highlight

In this paper, we present a new generic Structure-from-Motion pipeline, GenSfM, that uses a non-parametric camera projection model. The model is self-calibrated during the reconstruction process and can fit a wide variety of cameras, ranging from simple low-distortion pinhole cameras to more extreme…

2025

SuperDec: 3D Scene Decomposition with Superquadrics Primitives

ICCV 2025poster

We present SuperDec, an approach for compact 3D scene representations based on geometric primitives, namely superquadrics.While most recent works leverage geometric primitives to obtain photorealistic 3D scene representations, we propose to leverage them to obtain a compact yet expressive representa…

Cited by 0SourcePDFScholar
2025

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

CVPR 2025poster

Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundamentally does not scale well. We argue that learning from in-the-wild human videos…

Cited by 0SourcePDFScholar
2025

Video Perception Models for 3D Scene Synthesis

NeurIPS 2025poster

Automating the expert-dependent and labor-intensive task of 3D scene synthesis would significantly benefit fields such as architectural design, robotics simulation, and virtual reality. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or…

Cited by 0SourceScholar
2025

WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments

CVPR 2025poster

We present WildGS-SLAM, a robust and efficient monocular RGB SLAM system designed to handle dynamic environments by leveraging uncertainty-aware geometric mapping. Unlike traditional SLAM systems, which assume static scenes, our approach integrates depth and uncertainty information to enhance tracki…

2024

Active Visual Localization for Multi-Agent Collaboration: A Data-Driven Approach

ICRA 2024poster

Rather than having each newly deployed robot create its own map of its surroundings, the growing availability of SLAM-enabled devices provides the option of simply localizing in a map of another robot or device. In cases such as multi-robot or human-robot collaboration, localizing all agents in the…

Cited by 6SourceScholar
2024

CR3DT: Camera-RADAR Fusion for 3D Detection and Tracking

IROS 2024poster

To enable self-driving vehicles accurate detection and tracking of surrounding objects is essential. While Light Detection and Ranging (LiDAR) sensors have set the benchmark for high-performance systems, the appeal of camera-only solutions lies in their cost-effectiveness. Notably, despite the preva…

Cited by 11SourcecodeScholar
2024

Diffusion Bridges for 3D Point Cloud Denoising

ECCV 2024poster

"In this work, we address the task of point cloud denoising using a novel framework adapting Diffusion Schrödinger bridges to unstructured data like point sets. Unlike previous works that predict point-wise displacements from point features or learned noise distributions, our method learns an optim…

2024

Dynamic 3D Gaussian Fields for Urban Areas

NeurIPS 2024spotlight

We present an efficient neural 3D scene representation for novel-view synthesis (NVS) in large-scale, dynamic urban areas. Existing works are not well suited for applications like mixed-reality or closed-loop simulation due to their limited visual quality and non-interactive rendering speeds. Recent…

Cited by 14SourcePDFScholar
2024

EgoGen: An Egocentric Synthetic Data Generator

CVPR 2024poster

Understanding the world in first-person view is fundamental in Augmented Reality (AR). This immersive perspective brings dramatic visual changes and unique challenges compared to third-person views. Synthetic data has empowered third-person-view vision models but its application to embodied egocentr…

Cited by 17SourcePDFScholar
2024

F3Loc: Fusion and Filtering for Floorplan Localization

CVPR 2024highlight

In this paper we propose an efficient data-driven solution to self-localization within a floorplan. Floorplan data is readily available long-term persistent and inherently robust to changes in the visual appearance. Our method does not require retraining per map and location or demand a large databa…

Cited by 7SourcePDFScholar
2024

GLACE: Global Local Accelerated Coordinate Encoding

CVPR 2024poster

Scene coordinate regression (SCR) methods are a family of visual localization methods that directly regress 2D-3D matches for camera pose estimation. They are effective in small-scale scenes but face significant challenges in large-scale scenes that are further amplified in the absence of ground tru…

2024

GeneAvatar: Generic Expression-Aware Volumetric Head Avatar Editing from a Single Image

CVPR 2024poster

Recently we have witnessed the explosive growth of various volumetric representations in modeling animatable head avatars. However due to the diversity of frameworks there is no practical method to support high-level applications like 3D head avatar editing across different representations. In this…

2024

GeoCalib: Learning Single-image Calibration with Geometric Optimization

ECCV 2024poster

"From a single image, visual cues can help deduce intrinsic and extrinsic camera parameters like the focal length and the gravity direction. This single-image calibration can benefit various downstream applications like image editing and 3D mapping. Current approaches to this problem are based on ei…

2024

Global Structure-from-Motion Revisited

ECCV 2024poster

"Recovering 3D structure and camera motion from images has been a long-standing focus of computer vision research and is known as Structure-from-Motion (SfM). Solutions to this problem are categorized into incremental and global approaches. Until now, the most popular systems follow the incremental…

2024

Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning

CVPR 2024poster

Recovering the 3D scene geometry from a single view is a fundamental yet ill-posed problem in computer vision. While classical depth estimation methods infer only a 2.5D scene representation limited to the image plane recent approaches based on radiance fields reconstruct a full 3D representation. H…

2024

LEAP-VO: Long-term Effective Any Point Tracking for Visual Odometry

CVPR 2024poster

Visual odometry estimates the motion of a moving camera based on visual input. Existing methods mostly focusing on two-view point tracking often ignore the rich temporal context in the image sequence thereby overlooking the global motion patterns and providing no assessment of the full trajectory re…

2024

Learning Where to Look: Self-supervised Viewpoint Selection for Active Localization using Geometrical Information

ECCV 2024poster

"Accurate localization in diverse environments is a fundamental challenge in computer vision and robotics. The task involves determining a sensor’s precise position and orientation, typically a camera, within a given space. Traditional localization methods often rely on passive sensing, which may st…

2024

Leveraging Neural Radiance Fields for Uncertainty-Aware Visual Localization

ICRA 2024poster

As a promising fashion for visual localization, scene coordinate regression (SCR) has seen tremendous progress in the past decade. Most recent methods usually adopt neural networks to learn the mapping from image pixels to 3D scene coordinates, which requires a vast amount of annotated training data…

Cited by 12SourceScholar
2024

MAP-ADAPT: Real-Time Quality-Adaptive Semantic 3D Maps

ECCV 2024poster

"Creating 3D semantic reconstructions of environments is fundamental to many applications, especially when related to autonomous agent operation (, goal-oriented navigation or object interaction and manipulation). Commonly, 3D semantic reconstruction systems capture the entire scene in the same leve…

2024

MICDrop: Masking Image and Depth Features via Complementary Dropout for Domain-Adaptive Semantic Segmentation

ECCV 2024poster

"Unsupervised Domain Adaptation (UDA) is the task of bridging the domain gap between a labeled source domain, e.g., synthetic data, and an unlabeled target domain. We observe that current UDA methods show inferior results on fine structures and tend to oversegment objects with ambiguous appearance.…

2024

MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

ECCV 2024oral

"We introduce , an efficient model that, given sparse multi-view images as input, predicts clean feed-forward 3D Gaussians. To accurately localize the Gaussian centers, we build a cost volume representation via plane sweeping, where the cross-view feature similarities stored in the cost volume can p…

2024

MaRINeR: Enhancing Novel Views by Matching Rendered Images with Nearby References

ECCV 2024poster

"Rendering realistic images from 3D reconstruction is an essential task of many Computer Vision and Robotics pipelines, notably for mixed-reality applications as well as training autonomous agents in simulated environments. However, the quality of novel views heavily depends of the source reconstruc…

2024

MuRF: Multi-Baseline Radiance Fields

CVPR 2024poster

We present Multi-Baseline Radiance Fields (MuRF) a general feed-forward approach to solving sparse view synthesis under multiple different baseline settings (small and large baselines and different number of input views). To render a target novel view we discretize the 3D space into planes parallel…

2024

Multi-Level Neural Scene Graphs for Dynamic Urban Environments

CVPR 2024poster

We estimate the radiance field of large-scale dynamic areas from multiple vehicle captures under varying environmental conditions. Previous works in this domain are either restricted to static environments do not scale to more than a single short video or struggle to separately represent dynamic obj…

Cited by 10SourcePDFScholar
2024

Multiway Point Cloud Mosaicking with Diffusion and Global Optimization

CVPR 2024poster

We introduce a novel framework for multiway point cloud mosaicking (named Wednesday) designed to co-align sets of partially overlapping point clouds -- typically obtained from 3D scanners or moving RGB-D cameras -- into a unified coordinate system. At the core of our approach is ODIN a learned pairw…

2024

NeRF On-the-go: Exploiting Uncertainty for Distractor-free NeRFs in the Wild

CVPR 2024poster

Neural Radiance Fields (NeRFs) have shown remarkable success in synthesizing photorealistic views from multi-view images of static scenes but face challenges in dynamic real-world environments with distractors like moving objects shadows and lighting changes. Existing methods manage controlled envir…

2024

PickScan: Object discovery and reconstruction from handheld interactions

IROS 2024poster

Reconstructing compositional 3D representations of scenes, where each object is represented with its own 3D model, is a highly desirable capability in robotics and augmented reality. However, most existing methods rely heavily on strong appearance priors for object discovery, therefore only working…

Cited by 0SourcecodeScholar
2024

ResFields: Residual Neural Fields for Spatiotemporal Signals

ICLR 2024spotlight

Neural fields, a category of neural networks trained to represent high-frequency signals, have gained significant attention in recent years due to their impressive performance in modeling complex 3D data, such as signed distance (SDFs) or radiance fields (NeRFs), via a single multi-layer perceptron…

2024

Robust Incremental Structure-from-Motion with Hybrid Features

ECCV 2024poster

"Structure-from-Motion (SfM) has become a ubiquitous tool for camera calibration and scene reconstruction with many downstream applications in computer vision and beyond. While the state-of-the-art SfM pipelines have reached a high level of maturity in well-textured and well-configured scenes over t…

2024

SNI-SLAM: Semantic Neural Implicit SLAM

CVPR 2024poster

We propose SNI-SLAM a semantic SLAM system utilizing neural implicit representation that simultaneously performs accurate semantic mapping high-quality surface reconstruction and robust camera tracking. In this system we introduce hierarchical semantic representation to allow multi-level semantic co…

2024

SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D Scenes

CVPR 2024poster

Existing 3D scene understanding methods are heavily focused on 3D semantic and instance segmentation. However identifying objects and their parts only constitutes an intermediate step towards a more fine-grained goal which is effectively interacting with the functional interactive elements (e.g. han…

Cited by 35SourcePDFScholar
2024

SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs

ECCV 2024poster

"We introduce the task of localizing an input image within a multi-modal reference map represented by a collection of 3D scene graphs. These scene graphs comprise multiple modalities, including object-level point clouds, images, attributes, and relationships between objects, offering a lightweight a…

2024

Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels

ECCV 2024poster

"Current 3D scene segmentation methods are heavily dependent on manually annotated 3D training datasets. Such manual annotations are labor-intensive, and often lack fine-grained details. Furthermore, models trained on this data typically struggle to recognize object classes beyond the annotated trai…

Cited by 34SourcePDFScholar
2024

Semicalibrated Relative Pose from an Affine Correspondence and Monodepth

ECCV 2024poster

"We address the semi-calibrated relative pose estimation problem where we assume the principal point to be located in the center of the image and estimate the focal lengths, relative rotation, and translation of two cameras. We introduce the first minimal solver that requires only a single affine co…

2024

Spherical Frustum Sparse Convolution Network for LiDAR Point Cloud Semantic Segmentation

NeurIPS 2024poster

LiDAR point cloud semantic segmentation enables the robots to obtain fine-grained semantic information of the surrounding environment. Recently, many works project the point cloud onto the 2D image and adopt the 2D Convolutional Neural Networks (CNNs) or vision transformer for LiDAR point cloud sema…

2024

StereoGlue: Joint Feature Matching and Robust Estimation

ECCV 2024poster

"We propose StereoGlue, a method designed for joint feature matching and robust estimation that effectively reduces the combinatorial complexity of these tasks using single-point minimal solvers. StereoGlue is applicable to a range of problems, including but not limited to relative pose and homograp…

2024

UniSDF: Unifying Neural Representations for High-Fidelity 3D Reconstruction of Complex Scenes with Reflections

NeurIPS 2024poster

Neural 3D scene representations have shown great potential for 3D reconstruction from 2D images. However, reconstructing real-world captures of complex scenes still remains a challenge. Existing generic 3D reconstruction methods often struggle to represent fine geometric details and do not adequatel…

2024

WildGaussians: 3D Gaussian Splatting In the Wild

NeurIPS 2024poster

While the field of 3D scene reconstruction is dominated by NeRFs due to their photorealistic quality, 3D Gaussian Splatting (3DGS) has recently emerged, offering similar quality with real-time rendering speeds. However, both methods primarily excel with well-controlled 3D scenes, while in-the-wild d…

2024

WorldPose: A World Cup Dataset for Global 3D Human Pose Estimation

ECCV 2024poster

"We present , a novel dataset for advancing research in multi-person global pose estimation in the wild, featuring footage from the 2022 FIFA World Cup. While previous datasets have primarily focused on local poses, often limited to a single person or in constrained, indoor settings, the infrastruct…

Cited by 5SourcePDFScholar
2023

DeepLSD: Line Segment Detection and Refinement With Deep Image Gradients

CVPR 2023poster

Line segments are ubiquitous in our human-made world and are increasingly used in vision tasks. They are complementary to feature points thanks to their spatial extent and the structural information they provide. Traditional line detectors based on the image gradient are extremely fast and accurate,…

2023

Four-View Geometry With Unknown Radial Distortion

CVPR 2023poster

We present novel solutions to previously unsolved problems of relative pose estimation from images whose calibration parameters, namely focal lengths and radial distortion, are unknown. Our approach enables metric reconstruction without modeling these parameters. The minimal case for reconstruction…

Cited by 10SourcePDFScholar
2023

GlueStick: Robust Image Matching by Sticking Points and Lines Together

ICCV 2023poster

Line segments are powerful features complementary to points. They offer structural cues, robust to drastic viewpoint and illumination changes, and can be present even in texture-less areas. However, describing and matching them is more challenging compared to points due to partial occlusions, lack o…

Cited by 74PDFcodeScholar
2023

HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World

ICCV 2023poster

Building an interactive AI assistant that can perceive, reason, and collaborate with humans in the real world has been a long-standing pursuit in the AI community. This work is part of a broader research effort to develop intelligent agents that can interactively guide humans through performing task…

Cited by 55PDFcodeScholar
2023

Human from Blur: Human Pose Tracking from Blurry Images

ICCV 2023poster

We propose a method to estimate 3D human poses from substantially blurred images. The key idea is to tackle the inverse problem of image deblurring by modeling the forward problem with a 3D human model, a texture map, and a sequence of poses to describe human motion. The blurring process is then mod…

Cited by 3PDFScholar
2023

IntrinsicNeRF: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis

ICCV 2023poster

Existing inverse rendering combined with neural rendering methods can only perform editable novel view synthesis on object-specific scenes, while we present intrinsic neural radiance fields, dubbed IntrinsicNeRF, which introduce intrinsic decomposition into the NeRF-based neural rendering method and…

Cited by 58PDFcodeScholar
2023

Learning-Based Dimensionality Reduction for Computing Compact and Effective Local Feature Descriptors

ICRA 2023poster

A distinctive representation of image patches in form of features is a key component of many computer vision and robotics tasks, such as image matching, image retrieval, and visual localization. State-of-the-art descriptors, from hand-crafted descriptors such as SIFT to learned ones such as HardNet,…

Cited by 11SourcecodeScholar
2023

Learning-based Relational Object Matching Across Views

ICRA 2023poster

Intelligent robots require object-level scene understanding to reason about possible tasks and interactions with the environment. Moreover, many perception tasks such as scene reconstruction, image retrieval, or place recognition can benefit from reasoning on the level of objects. While keypoint-bas…

Cited by 5SourceScholar
2023

OpenMask3D: Open-Vocabulary 3D Instance Segmentation

NeurIPS 2023poster

We introduce the task of open-vocabulary 3D instance segmentation. Current approaches for 3D instance segmentation can typically only recognize object categories from a pre-defined closed set of classes that are annotated in the training datasets. This results in important limitations for real-world…

2023

OpenScene: 3D Scene Understanding With Open Vocabularies

CVPR 2023poster

Traditional 3D scene understanding approaches rely on labeled 3D datasets to train a model for a single task with supervision. We propose OpenScene, an alternative approach where a model predicts dense features for 3D scene points that are co-embedded with text and image pixels in CLIP feature space…

Cited by 341SourcePDFScholar
2023

Privacy Preserving Localization via Coordinate Permutations

ICCV 2023poster

Recent methods on privacy-preserving image-based localization use a random line parameterization to protect the privacy of query images and database maps. The lifting of points to lines effectively drops one of the two geometric constraints traditionally used with point-to-point correspondences in s…

Cited by 4PDFScholar
2023

R3D3: Dense 3D Reconstruction of Dynamic Scenes from Multiple Cameras

ICCV 2023poster

Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera systems provide a simpler, low-cost alternative. However, camera-based 3D reconstruction of complex dynamic scenes has pr…

Cited by 30PDFScholar
2023

RLSAC: Reinforcement Learning Enhanced Sample Consensus for End-to-End Robust Estimation

ICCV 2023poster

Robust estimation is a crucial and still challenging task, which involves estimating model parameters in noisy environments. Although conventional sampling consensus-based algorithms sample several times to achieve robustness, these algorithms cannot use data features and historical information effe…

Cited by 6PDFcodeScholar
2023

RegFormer: An Efficient Projection-Aware Transformer Network for Large-Scale Point Cloud Registration

ICCV 2023poster

Although point cloud registration has achieved remarkable advances in object-level and indoor scenes, large-scale registration methods are rarely explored. Challenges mainly arise from the huge point number, complex distribution, and outliers of outdoor LiDAR scans. In addition, most existing regist…

Cited by 58PDFcodeScholar
2023

Removing Objects From Neural Radiance Fields

CVPR 2023poster

Neural Radiance Fields (NeRFs) are emerging as a ubiquitous scene representation that allows for novel view synthesis. Increasingly, NeRFs will be shareable with other people. Before sharing a NeRF, though, it might be desirable to remove personal information or unsightly objects. Such removal is no…

Cited by 71SourcePDFScholar
2023

SGAligner: 3D Scene Alignment with Scene Graphs

ICCV 2023poster

Building 3D scene graphs has recently emerged as a topic in scene representation for several embodied AI applications to represent the world in a structured and rich manner. With their increased use in solving downstream tasks (e.g., navigation and room rearrangement), can we leverage and recycle th…

Cited by 14PDFcodeScholar
2023

SNAP: Self-Supervised Neural Maps for Visual Positioning and Semantic Understanding

NeurIPS 2023poster

Semantic 2D maps are commonly used by humans and machines for navigation purposes, whether it's walking or driving. However, these maps have limitations: they lack detail, often contain inaccuracies, and are difficult to create and maintain, especially in an automated fashion. Can we use _raw image…

2023

The Drunkard’s Odometry: Estimating Camera Motion in Deforming Scenes

NeurIPS 2023poster

Estimating camera motion in deformable scenes poses a complex and open research challenge. Most existing non-rigid structure from motion techniques assume to observe also static scene parts besides deforming scene parts in order to establish an anchoring reference. However, this assumption does not…

2023

Tracking by 3D Model Estimation of Unknown Objects in Videos

ICCV 2023poster

Most model-free visual object tracking methods formulate the tracking task as object location estimation given by a 2D segmentation or a bounding box in each video frame. We argue that this representation is limited and instead propose to guide and improve 2D tracking with an explicit object represe…

Cited by 7PDFScholar
2023

Vanishing Point Estimation in Uncalibrated Images with Prior Gravity Direction

ICCV 2023poster

We tackle the problem of estimating a Manhattan frame, i.e. three orthogonal vanishing points, and the unknown focal length of the camera, leveraging a prior vertical direction. The direction can come from an Inertial Measurement Unit that is a standard component of recent consumer devices, e.g., sm…

Cited by 4PDFcodeScholar
2023

VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View Reconstruction

CVPR 2023poster

The success of the Neural Radiance Fields (NeRF) in novel view synthesis has inspired researchers to propose neural implicit scene reconstruction. However, most existing neural implicit reconstruction methods optimize per-scene parameters and therefore lack generalizability to new scenes. We introdu…

2022

CompNVS: Novel View Synthesis with Scene Completion

ECCV 2022poster

"We introduce a scalable framework for novel view synthesis from RGB-D images with largely incomplete scene coverage. While generative neural approaches have demonstrated spectacular results on 2D images, they have not yet achieved similar photorealistic results in combination with scene completion…

Cited by 8SourcePDFScholar
2022

EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices

ECCV 2022poster

"Understanding social interactions from egocentric views is crucial for many applications, ranging from assistive robotics to AR/VR. Key to reasoning about interactions is to understand the body pose and motion of the interaction partner from the egocentric view. However, research in this area is se…

2022

IterMVS: Iterative Probability Estimation for Efficient Multi-View Stereo

CVPR 2022poster

We present IterMVS, a new data-driven method for high-resolution multi-view stereo. We propose a novel GRU-based estimator that encodes pixel-wise probability distributions of depth in its hidden state. Ingesting multi-scale matching information, our model refines these distributions over multiple i…

Cited by 123PDFcodeScholar
2022

LaMAR: Benchmarking Localization and Mapping for Augmented Reality

ECCV 2022poster

"Localization and mapping is the foundational technology for augmented reality (AR) that enables sharing and persistence of digital content in the real world. While significant progress has been made, researchers are still mostly driven by unrealistic benchmarks not representative of real-world AR s…

2022

Learning To Align Sequential Actions in the Wild

CVPR 2022poster

State-of-the-art methods for self-supervised sequential action alignment rely on deep networks that find correspondences across videos in time. They either learn frame-to-frame mapping across sequences, which does not leverage temporal information, or assume monotonic alignment between each video pa…

Cited by 32PDFcodeScholar
2022

Motion-From-Blur: 3D Shape and Motion Estimation of Motion-Blurred Objects in Videos

CVPR 2022poster

We propose a method for jointly estimating the 3D motion, 3D shape, and appearance of highly motion-blurred objects from a video. To this end, we model the blurred appearance of a fast moving object in a generative fashion by parametrizing its 3D position, rotation, velocity, acceleration, bounces,…

Cited by 10PDFcodeScholar
2022

NICE-SLAM: Neural Implicit Scalable Encoding for SLAM

CVPR 2022poster

Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over-smoothed scene reconstructions and have difficulty scaling up to large scenes. These l…

Cited by 766PDFcodeScholar
2022

Panoptic Multi-TSDFs: a Flexible Representation for Online Multi-resolution Volumetric Mapping and Long-term Dynamic Scene Consistency

ICRA 2022poster

For robotic interaction in environments shared with other agents, access to volumetric and semantic maps of the scene is crucial. However, such environments are inevitably subject to long-term changes, which the map needs to account for. We thus propose panoptic multi-TSDFs as a novel representation…

Cited by 74SourcecodeScholar
2021

Back to the Feature: Learning Robust Camera Localization From Pixels To Pose

CVPR 2021poster

Camera pose estimation in known scenes is a 3D geometry task recently tackled by multiple learning algorithms. Many regress precise geometric quantities, like poses or 3D points, from an input image. This either fails to generalize to new viewpoints or ties the model parameters to a specific scene.…

Cited by 301PDFcodeScholar
2021

CodeVIO: Visual-Inertial Odometry with Learned Optimizable Dense Depth

ICRA 2021poster

In this work, we present a lightweight, tightly-coupled deep depth network and visual-inertial odometry (VIO) system, which can provide accurate state estimates and dense depth maps of the immediate surroundings. Leveraging the proposed lightweight Conditional Variational Autoencoder (CVAE) for dept…

Cited by 52SourceScholar
2021

Cross-Descriptor Visual Localization and Mapping

ICCV 2021poster

Visual localization and mapping is the key technology underlying the majority of mixed reality and robotics systems. Most state-of-the-art approaches rely on local features to establish correspondences between images. In this paper, we present three novel scenarios for localization and mapping which…

Cited by 35PDFcodeScholar
2021

DeFMO: Deblurring and Shape Recovery of Fast Moving Objects

CVPR 2021poster

Objects moving at high speed appear significantly blurred when captured with cameras. The blurry appearance is especially ambiguous when the object has complex shape or texture. In such cases, classical methods, or even humans, are unable to recover the object's appearance and motion. We propose a m…

Cited by 50PDFcodeScholar
2021

DeepVideoMVS: Multi-View Stereo on Video With Recurrent Spatio-Temporal Fusion

CVPR 2021poster

We propose an online multi-view depth prediction approach on posed video streams, where the scene geometry information computed in the previous time steps is propagated to the current time step in an efficient and geometrically plausible way. The backbone of our approach is a real-time capable, ligh…

Cited by 118PDFcodeScholar
2021

FMODetect: Robust Detection of Fast Moving Objects

ICCV 2021poster

We propose the first learning-based approach for fast moving objects detection. Such objects are highly blurred and move over large distances within one video frame. Fast moving objects are associated with a deblurring and matting problem, also called deblatting. We show that the separation of debla…

Cited by 14PDFcodeScholar
2021

H2O: Two Hands Manipulating Objects for First Person Interaction Recognition

ICCV 2021poster

We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects. To this end, we propose a method to create a unified dataset for egocentric 3D interaction recognition. Our method produces annotations of the 3D pose of two…

Cited by 205PDFScholar
2021

Holistic 3D Scene Understanding From a Single Image With Implicit Representation

CVPR 2021poster

We present a new pipeline for holistic 3D scene understanding from a single image, which could predict object shape, object pose and scene layout. As it is a highly ill-posed problem, existing methods usually suffer from inaccurate estimation of both shapes and layout especially for the cluttered sc…

Cited by 129PDFcodeScholar
2021

Learning Motion Priors for 4D Human Body Capture in 3D Scenes

ICCV 2021poster

Recovering high-quality 3D human motion in complex scenes from monocular videos is important for many applications, ranging from AR/VR to robotics. However, capturing realistic human-scene interactions, while dealing with occlusions and partial views, is challenging; current approaches are still far…

Cited by 116PDFcodeScholar
2021

NeuralFusion: Online Depth Fusion in Latent Space

CVPR 2021poster

We present a novel online depth map fusion approach that learns depth map aggregation in a latent feature space. While previous fusion methods use an explicit scene representation like signed distance functions (SDFs), we propose a learned feature representation for the fusion. The key idea is a sep…

Cited by 65PDFcodeScholar
2021

PatchmatchNet: Learned Multi-View Patchmatch Stereo

CVPR 2021poster

We present PatchmatchNet, a novel and learnable cascade formulation of Patchmatch for high-resolution multi-view stereo. With high computation speed and low memory requirement, PatchmatchNet can process higher resolution imagery and is more suited to run on resource limited devices than competitors…

Cited by 391PDFcodeScholar
2021

Pixel-Perfect Structure-From-Motion With Featuremetric Refinement

ICCV 2021poster

Finding local features that are repeatable across multiple views is a cornerstone of sparse 3D reconstruction. The classical image matching paradigm detects keypoints per-image once and for all, which can yield poorly-localized features and propagate large errors to the final geometry. In this paper…

Cited by 209PDFcodeScholar
2021

Privacy Preserving Localization and Mapping From Uncalibrated Cameras

CVPR 2021poster

Recent works on localization and mapping from privacy preserving line features have made significant progress towards addressing the privacy concerns arising from cloud-based solutions in mixed reality and robotics. The requirement for calibrated cameras is a fundamental limitation for these approac…

Cited by 15PDFScholar
2021

Privacy-Preserving Image Features via Adversarial Affine Subspace Embeddings

CVPR 2021poster

Many computer vision systems require users to upload image features to the cloud for processing and storage. These features can be exploited to recover sensitive information about the scene or subjects, e.g., by reconstructing the appearance of the original image. To address this privacy concern, we…

Cited by 41PDFScholar
2021

SOLD2: Self-Supervised Occlusion-Aware Line Description and Detection

CVPR 2021poster

Compared to feature point detection and description, detecting and matching line segments offer additional challenges. Yet, line features represent a promising complement to points for multi-view tasks. Lines are indeed well-defined by the image gradient, frequently appear even in poorly textured ar…

Cited by 97PDFcodeScholar
2021

Sat2Vid: Street-View Panoramic Video Synthesis From a Single Satellite Image

ICCV 2021poster

We present a novel method for synthesizing both temporally and geometrically consistent street-view panoramic video from a single satellite image and camera trajectory. Existing cross-view synthesis approaches focus on images, while video synthesis in such a case has not yet received enough attentio…

Cited by 12PDFScholar
2021

Shape As Points: A Differentiable Poisson Solver

NeurIPS 2021oral

In recent years, neural implicit representations gained popularity in 3D reconstruction due to their expressiveness and flexibility. However, the implicit nature of neural implicit representations results in slow inference times and requires careful initialization. In this paper, we revisit the clas…

2021

Shape from Blur: Recovering Textured 3D Shape and Motion of Fast Moving Objects

NeurIPS 2021poster

We address the novel task of jointly reconstructing the 3D shape, texture, and motion of an object from a single motion-blurred image. While previous approaches address the deblurring problem only in the 2D image domain, our proposed rigorous modeling of all object properties in the 3D domain enable…

2021

Towards Efficient Graph Convolutional Networks for Point Cloud Handling

ICCV 2021poster

We aim at improving the computational efficiency of graph convolutional networks (GCNs) for learning on point clouds. The basic graph convolution that is composed of a K-nearest neighbor (KNN) search and a multilayer perceptron (MLP) is examined. By mathematically analyzing the operations there, two…

Cited by 34PDFScholar
2020

Aerial Single-View Depth Completion With Image-Guided Uncertainty Estimation

RA-L 2020

On the pursuit of autonomous flying robots, the scientific community has been developing onboard real-time algorithms for localisation, mapping and planning. Despite recent progress, the available solutions still lack accuracy and robustness in many aspects. While mapping for autonomous cars had a s

Cited by 60SourcecodeScholar
2020

Calibration-free Structure-from-Motion with Calibrated Radial Trifocal Tensors

ECCV 2020poster

In this paper we consider the problem of Structure-from-Motion from images with unknown intrinsic calibration. Instead of estimating the internal camera parameters through some self-calibration procedure, we propose to use a subset of the reprojection constraints that is invariant to radial displace…

Cited by 20SourcePDFScholar
2020

Convolutional Occupancy Networks

ECCV 2020poster

Recently, implicit neural representations have gained popularity for learning-based 3D reconstruction. While demonstrating promising results, most implicit approaches are limited to comparably simple geometry of single objects and do not scale to more complicated or large-scale scenes. The key limit…

2020

DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere Tracing

CVPR 2020poster

We propose a differentiable sphere tracing algorithm to bridge the gap between inverse graphics methods and the recently proposed deep learning based implicit signed distance function. Due to the nature of the implicit function, the rendering process requires tremendous function queries, which is pa…

Cited by 350PDFcodeScholar
2020

Geometry-Aware Satellite-to-Ground Image Synthesis for Urban Areas

CVPR 2020poster

We present a novel method for generating panoramic street-view images which are geometrically consistent with a given satellite image. Different from existing approaches that completely rely on a deep learning architecture to generalize cross-view image distributions, our approach explicitly loops i…

Cited by 78PDFScholar
2020

Handcrafted Outlier Detection Revisited

ECCV 2020poster

Local feature matching is a critical part of many computer vision pipelines, including among others Structure-from-Motion, SLAM, and Visual Localization. However, due to limitations in the descriptors, raw matches are often contaminated by a majority of outliers. As a result, outlier detection is a…

2020

Infrastructure-based Multi-Camera Calibration using Radial Projections

ECCV 2020poster

Multi-camera systems are an important sensor platform for intelligent systems such as self-driving cars. Pattern-based calibration techniques can be used to calibrate the intrinsics of the cameras individually. However, extrinsic calibration of systems with little to no visual overlap between the ca…

2020

LIC-Fusion 2.0: LiDAR-Inertial-Camera Odometry with Sliding-Window Plane-Feature Tracking

IROS 2020poster

Multi-sensor fusion of multi-modal measurements from commodity inertial, visual and LiDAR sensors to provide robust and accurate 6DOF pose estimation holds great potential in robotics and beyond. In this paper, building upon our prior work (i.e., LIC-Fusion), we develop a sliding-window filter based…

Cited by 150SourceScholar
2020

Leveraging Photometric Consistency Over Time for Sparsely Supervised Hand-Object Reconstruction

CVPR 2020poster

Modeling hand-object manipulations is essential for understanding how humans interact with their environment. While of practical importance, estimating the pose of hands and objects during interactions is challenging due to the large mutual occlusions that occur during manipulation. Recent efforts h…

Cited by 217PDFScholar
2020

Multi-View Optimization of Local Feature Geometry

ECCV 2020poster

In this work, we address the problem of refining the geometry of local image features from multiple views without known scene or camera geometry. Current approaches to local feature detection are inherently limited in their keypoint localization accuracy because they only operate on a single view. T…

2020

OmniSLAM: Omnidirectional Localization and Dense Mapping for Wide-baseline Multi-camera Systems

ICRA 2020poster

In this paper, we present an omnidirectional localization and dense mapping system for a wide-baseline multiview stereo setup with ultra-wide field-of-view (FOV) fisheye cameras, which has a 360° coverage of stereo observations of the environment. For more practical and accurate reconstruction, we f…

Cited by 60SourceScholar
2020

Online Invariance Selection for Local Feature Descriptors

ECCV 2020poster

To be invariant, or not to be invariant: that is the question formulated in this work about local descriptors. A limitation of current feature descriptors is the trade-off between generalization and discriminative power: more invariance means less informative descriptors. We propose to overcome this…

2020

Privacy Preserving Structure-from-Motion

ECCV 2020poster

Over the last years, visual localization and mapping solutions have been adopted by an increasing number of mixed reality and robotics systems. The recent trend towards cloud-based localization and mapping systems has raised significant privacy concerns. These are mainly grounded by the fact that th…

Cited by 48SourcePDFScholar
2020

RoutedFusion: Learning Real-Time Depth Map Fusion

CVPR 2020oral

The efficient fusion of depth maps is a key part of most state-of-the-art 3D reconstruction methods. Besides requiring high accuracy, these depth fusion methods need to be scalable and real-time capable. To this end, we present a novel real-time capable machine learning-based method for depth map fu…

Cited by 97PDFcodeScholar
2020

Self-Supervised Human Depth Estimation From Monocular Videos

CVPR 2020poster

Previous methods on estimating detailed human depth often require supervised training with 'ground truth' depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes training data collection simple and improves the generalization of…

Cited by 35PDFScholar
2020

To Learn or Not to Learn: Visual Localization from Essential Matrices

ICRA 2020poster

Visual localization is the problem of estimating a camera within a scene and a key technology for autonomous robots. State-of-the-art approaches for accurate visual localization use scene-specific representations, resulting in the overhead of constructing these models when applying the techniques to…

Cited by 128SourceScholar
2020

Why Having 10,000 Parameters in Your Camera Model Is Better Than Twelve

CVPR 2020oral

Camera calibration is an essential first step in setting up 3D Computer Vision systems. Commonly used parametric camera models are limited to a few degrees of freedom and thus often do not optimally fit to complex real lens distortion. In contrast, generic camera models allow for very accurate calib…

Cited by 70PDFcodeScholar
2019

3D Appearance Super-Resolution With Deep Learning

CVPR 2019poster

We tackle the problem of retrieving high-resolution (HR) texture maps of objects that are captured from multiple view points. In the multi-view case, model-based super-resolution (SR) methods have been recently proved to recover high quality texture maps. On the other hand, the advent of deep learn…

Cited by 41PDFcodeScholar
2019

A Cross-Season Correspondence Dataset for Robust Semantic Segmentation

CVPR 2019poster

In this paper, we present a method to utilize 2D-2D point matches between images taken during different image conditions to train a convolutional neural network for semantic segmentation. Enforcing label consistency across the matches makes the final segmentation algorithm robust to seasonal changes…

Cited by 104PDFcodeScholar
2019

D2-Net: A Trainable CNN for Joint Description and Detection of Local Features

CVPR 2019poster

In this work we address the problem of finding reliable pixel-level correspondences under difficult imaging conditions. We propose an approach where a single convolutional neural network plays a dual role: It is simultaneously a dense feature descriptor and a feature detector. By postponing the dete…

Cited by 909PDFcodeScholar
2019

DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene From Sparse LiDAR Data and Single Color Image

CVPR 2019poster

In this paper, we propose a deep learning architecture that produces accurate dense depth for the outdoor scene from a single color image and a sparse depth. Inspired by the indoor depth completion, our network estimates surface normals as the intermediate representation to produce dense depth, and…

Cited by 459PDFScholar
2019

Efficient 2D-3D Matching for Multi-Camera Visual Localization

ICRA 2019poster

Visual localization, i.e., determining the position and orientation of a vehicle with respect to a map, is a key problem in autonomous driving. We present a multi-camera visual inertial localization algorithm for large scale environments. To efficiently and effectively match features against a pre-b…

Cited by 41SourceScholar
2019

Episodic Curiosity through Reachability

ICLR 2019poster

Rewards are sparse in the real world and most of today's reinforcement learning algorithms struggle with such sparsity. One solution to this problem is to allow the agent to create rewards for itself - thus making rewards dense and more suitable for learning. In particular, inspired by curious behav…

2019

Incremental Visual-Inertial 3D Mesh Generation with Structural Regularities

ICRA 2019poster

Visual-Inertial Odometry (VIO) algorithms typically rely on a point cloud representation of the scene that does not model the topology of the environment. A 3D mesh instead offers a richer, yet lightweight, model. Nevertheless, building a 3D mesh out of the sparse and noisy 3D landmarks triangulated…

Cited by 64SourcecodeScholar
2019

Night-to-Day Image Translation for Retrieval-based Localization

ICRA 2019poster

Visual localization is a key step in many robotics pipelines, allowing the robot to (approximately) determine its position and orientation in the world. An efficient and scalable approach to visual localization is to use image retrieval techniques. These approaches identify the image most similar to…

Cited by 271SourcecodeScholar
2019

Privacy Preserving Image Queries for Camera Localization

ICCV 2019oral

Augmented/mixed reality and robotic applications are increasingly relying on cloud-based localization services, which require users to upload query images to perform camera pose estimation on a server. This raises significant privacy concerns when consumers use such services in their homes or in con…

Cited by 39PDFScholar
2019

Privacy Preserving Image-Based Localization

CVPR 2019poster

Image-based localization is a core component of many augmented/mixed reality (AR/MR) and autonomous robotic systems. Current localization systems rely on the persistent storage of 3D point clouds of the scene to enable camera pose estimation, but such data reveals potentially sensitive scene informa…

Cited by 99PDFScholar
2019

Project AutoVision: Localization and 3D Scene Perception for an Autonomous Vehicle with a Multi-Camera System

ICRA 2019poster

Project AutoVision aims to develop localization and 3D scene perception capabilities for a self-driving vehicle. Such capabilities will enable autonomous navigation in urban and rural environments, in day and night, and with cameras as the only exteroceptive sensors. The sensor suite employs many ca…

Cited by 149SourceScholar
2019

Real-Time Dense Mapping for Self-Driving Vehicles using Fisheye Cameras

ICRA 2019poster

We present a real-time dense geometric mapping algorithm for large-scale environments. Unlike existing methods which use pinhole cameras, our implementation is based on fisheye cameras whose large field of view benefits various computer vision applications for self-driving vehicles such as visual-in…

Cited by 47SourceScholar
2019

Reflection Separation using a Pair of Unpolarized and Polarized Images

NeurIPS 2019spotlight

When we take photos through glass windows or doors, the transmitted background scene is often blended with undesirable reflection. Separating two layers apart to enhance the image quality is of vital importance for both human and machine perception. In this paper, we propose to exploit physical cons…

2019

Understanding the Limitations of CNN-Based Absolute Camera Pose Regression

CVPR 2019poster

Visual localization is the task of accurate camera pose estimation in a known scene. It is a key problem in computer vision and robotics, with applications including self-driving cars, Structure-from-Motion, SLAM, and Mixed Reality. Traditionally, the localization problem has been tackled using 3D g…

Cited by 469PDFcodeScholar
2018

A Dataset of Flash and Ambient Illumination Pairs from the Crowd

ECCV 2018poster

Illumination is a critical element of photography and is essential for many computer vision tasks. Flash light is unique in the sense that it is a widely available tool for easily manipulating the scene illumination. We present a dataset of thousands of ambient and flash illumination pairs to enable…

Cited by 49SourcePDFScholar
2018

Augmenting Crowd-Sourced 3D Reconstructions Using Semantic Detections

CVPR 2018poster

Image-based 3D reconstruction for Internet photo collections has become a robust technology to produce impressive virtual representations of real-world scenes. However, several fundamental challenges remain for Structure-from-Motion (SfM) pipelines, namely: the placement and reconstruction of transi…

Cited by 9SourcePDFScholar
2018

Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions

CVPR 2018poster

Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing condition, including day-night changes, as well as weather and…

Cited by 780SourcePDFScholar
2018

Consensus Maximization for Semantic Region Correspondences

CVPR 2018poster

We propose a novel method for the geometric registration of semantically labeled regions. We approximate semantic regions by ellipsoids, and leverage their convexity to formulate the correspondence search effectively as a constrained optimization problem that maximizes the number of matched regions,…

Cited by 8SourcePDFScholar