← Search

Javier Civera

52 accepted papers

2026

Aion: Towards Hierarchical 4D Scene Graphs with Temporal Flow Dynamics

ICRA 2026poster

Autonomous navigation in dynamic environments requires spatial representations that capture both semantic structure and temporal evolution. 3D Scene Graphs (3DSGs) provide hierarchical multi-resolution abstractions that encode geometry and semantics, but existing extensions toward dynamics largely f…

2026

Neural Predictor-Corrector: Solving Homotopy Problems with Reinforcement Learning

ICLR 2026poster

The Homotopy paradigm, a general principle for solving challenging problems, appears across diverse domains such as robust optimization, global optimization, polynomial root-finding, and sampling. Practical solvers for these problems typically follow a predictor-corrector (PC) structure, but rely on…

Cited by 0SourceScholar
2026

S-Graphs 2.0 – a Hierarchical-Semantic Optimization and Loop Closure for SLAM

ICRA 2026poster

The hierarchical nature of 3D scene graphs aligns well with the structure of man-made environments, making them highly suitable for representation purposes. Beyond this, however, their embedded semantics and geometry could also be leveraged to improve the efficiency of map and pose optimization, an …

2026

Tightly Coupled SLAM with Imprecise Architectural Plans

ICRA 2026poster

Robots navigating indoor environments often have access to architectural plans, which can serve as prior knowledge to enhance their localization and mapping capabilities. While some SLAM algorithms leverage these plans for global localization in real-world environments, they typically overlook a cri…

2025

AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera Calibration

ICCV 2025poster

We present AnyCalib, a method for calibrating the intrinsic parameters of a camera from a single in-the-wild image, that is agnostic to the camera model. Current methods are predominantly tailored to specific camera models and/or require extrinsic cues, such as the direction of gravity, to be visibl…

2025

MRMT-PR: A Multi-Scale Reverse-View Mamba-Transformer for LiDAR Place Recognition

IROS 2025

Place recognition is a fundamental technology of high relevance for autonomous robot navigation. Existing methods encounter significant challenges arising from scene variations (e.g., illumination changes, dynamic objects), view-point shifts, and difficulties in data fusion and alignment. These fact

Cited by 1SourceScholar
2025

MVSAnywhere: Zero-Shot Multi-View Stereo

CVPR 2025poster

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision.However, most existing approaches do not generalize well across different domains and scene types (e.g. indoor vs outdoor). Training a general-purpose multi-view stereo model is challenging an…

2025

S-Graphs 2.0 - A Hierarchical-Semantic Optimization and Loop Closure for SLAM

RA-L 2025

The hierarchical nature of 3D scene graphs aligns well with the structure of man-made environments, making them highly suitable for representation purposes. Beyond this, however, their embedded semantics and geometry could also be leveraged to improve the efficiency of map and pose optimization, an

Cited by 7SourcecodeScholar
2025

Single-Shot Metric Depth from Focused Plenoptic Cameras

ICRA 2025

Metric depth estimation from visual sensors is crucial for robots to perceive, navigate, and interact with their environment. Traditional range imaging setups, such as stereo or structured light cameras, face hassles including calibration, occlusions, and hardware demands, with accuracy limited by t

Cited by 1SourceScholar
2025

Tightly Coupled SLAM With Imprecise Architectural Plans

RA-L 2025

Robots navigating indoor environments often have access to architectural plans, which can serve as prior knowledge to enhance their localization and mapping capabilities. While some SLAM algorithms leverage these plans for global localization in real-world environments, they typically overlook a cri

Cited by 4SourceScholar
2025

VSLAM-LAB: A Comprehensive Framework for Visual SLAM Methods and Datasets

IROS 2025

Visual Simultaneous Localization and Mapping (VSLAM) research faces significant challenges due to fragmented toolchains, complex system configurations, and inconsistent evaluation methodologies. To address these issues, we present VSLAM-LAB, a unified framework designed to streamline the development

Cited by 4SourcecodeScholar
2024

Adaptive Outlier Thresholding for Bundle Adjustment in Visual SLAM

ICRA 2024poster

State-of-the-art V-SLAM pipelines utilize robust cost functions and outlier rejection techniques to remove incorrect correspondences. However, these methods are typically fine-tuned to overfit certain benchmarks and struggle to adapt effectively to changes in the application domain or environmental…

Cited by 1SourcecodeScholar
2024

AnyFeature-VSLAM: Automating the Usage of Any Feature into Visual SLAM

RSS 2024poster

Feature-based SLAM heavily relies on the specific type of visual features employed. The most effective feature in some conditions may perform worse or not be suitable for other ones, leading to significant performance variability. Seamlessly switching to the most effective visual feature is a desira…

2024

From Correspondences to Pose: Non-minimal Certifiably Optimal Relative Pose without Disambiguation

CVPR 2024highlight

Estimating the relative camera pose from n \geq 5 correspondences between two calibrated views is a fundamental task in computer vision. This process typically involves two stages: 1) estimating the essential matrix between the views and 2) disambiguating among the four candidate relative poses that…

2024

Unifying Local and Global Multimodal Features for Place Recognition in Aliased and Low-Texture Environments

ICRA 2024poster

Perceptual aliasing and weak textures pose significant challenges to the task of place recognition, hindering the performance of Simultaneous Localization and Mapping (SLAM) systems. This paper presents a novel model, called UMF (standing for Unifying Local and Global Multimodal Features) that 1) le…

Cited by 3SourcecodeScholar
2023

Graph-Based Global Robot Localization Informing Situational Graphs with Architectural Graphs

IROS 2023poster

In this paper, we propose a solution for legged robot localization using architectural plans. Our specific contributions towards this goal are several. Firstly, we develop a method for converting the plan of a building into what we denote as an architectural graph (A-Graph). When the robot starts mo…

Cited by 11SourceScholar
2023

LightDepth: Single-View Depth Self-Supervision from Illumination Decline

ICCV 2023poster

Single-view depth estimation can be remarkably effective if there is enough ground-truth depth data for supervised training. However, there are scenarios, especially in medicine in the case of endoscopies, where such data cannot be obtained. In such cases, multi-view self-supervision and synthetic-t…

Cited by 8PDFScholar
2023

S-Graphs+: Real-Time Localization and Mapping Leveraging Hierarchical Representations

RA-L 2023

In this paper, we present an evolved version of Situational Graphs, which jointly models in a single optimizable factor graph (1) a pose graph, as a set of robot keyframes comprising associated measurements and robot poses, and (2) a 3D scene graph, as a high-level representation of the environment

Cited by 62SourcecodeScholar
2023

SID-SLAM: Semi-Direct Information-Driven RGB-D SLAM

RA-L 2023

This work presents SID-SLAM, a complete SLAM framework for RGB-D cameras. Our main contribution is a semi-direct approach that, for the first time, combines tightly and indistinctly photometric and feature-based image measurements. Additionally, SID-SLAM uses information metrics to reduce the state

Cited by 16SourceScholar
2023

SfM-TTR: Using Structure From Motion for Test-Time Refinement of Single-View Depth Networks

CVPR 2023poster

Estimating a dense depth map from a single view is geometrically ill-posed, and state-of-the-art methods rely on learning depth's relation with visual appearance using deep neural networks. On the other hand, Structure from Motion (SfM) leverages multi-view constraints to produce very accurate but s…

2023

The Drunkard’s Odometry: Estimating Camera Motion in Deforming Scenes

NeurIPS 2023poster

Estimating camera motion in deformable scenes poses a complex and open research challenge. Most existing non-rigid structure from motion techniques assume to observe also static scene parts besides deforming scene parts in order to establish an anchoring reference. However, this assumption does not…

2022

Bayesian Deep Neural Networks for Supervised Learning of Single-View Depth

RA-L 2022

Uncertainty quantification is essential for robotic perception, as overconfident or point estimators can lead to collisions and damages to the environment and the robot. In this letter, we evaluate scalable approaches to uncertainty quantification in single-view supervised depth learning, specifical

Cited by 10SourceScholar
2022

Danish Airs and Grounds: A Dataset for Aerial-to-Street-Level Place Recognition and Localization

RA-L 2022

Place recognition and visual localization are particularly challenging in wide baseline configurations. In this letter, we contribute with the <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Danish Airs and Grounds</i> (DAG) dataset, a large collecti

Cited by 12SourceScholar
2022

Jacobian Computation for Cumulative B-Splines on SE(3) and Application to Continuous-Time Object Tracking

RA-L 2022

In this paper we propose a method that estimates the <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$SE(3)$</tex-math></inline-formula> continuous trajectories (orientation and translation) of the dynamic rigid obj

Cited by 8SourceScholar
2022

Model for Multi-View Residual Covariances Based on Perspective Deformation

RA-L 2022

In this work, we derive a model for the covariance of the visual residuals in multi-view SfM, odometry and SLAM setups. The core of our approach is the formulation of the residual covariances as a combination of geometric and photometric noise sources. And our key novel contribution is the derivatio

Cited by 9SourceScholar
2022

Situational Graphs for Robot Navigation in Structured Indoor Environments

RA-L 2022

Mobile robots should be aware of their <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">situation</i> , comprising the deep understanding of their surrounding environment along with the estimation of its own state, to successfully make intelligent dec

Cited by 66SourceScholar
2021

Bayesian Triplet Loss: Uncertainty Quantification in Image Retrieval

ICCV 2021poster

Uncertainty quantification in image retrieval is crucial for downstream decisions, yet it remains a challenging and largely unexplored problem. Current methods for estimating uncertainties are poorly calibrated, computationally expensive, or based on heuristics. We present a new method that views im…

Cited by 39PDFScholar
2021

DOT: Dynamic Object Tracking for Visual SLAM

ICRA 2021poster

In this paper we present DOT (Dynamic Object Tracking), a front-end that added to existing SLAM systems can significantly improve their robustness and accuracy in highly dynamic environments. DOT combines instance segmentation and multi-view geometry to generate masks for dynamic objects in order to…

Cited by 91SourceScholar
2021

Endo-Depth-and-Motion: Reconstruction and Tracking in Endoscopic Videos Using Depth Networks and Photometric Constraints

RA-L 2021

Estimating a scene reconstruction and the camera motion from in-body videos is challenging due to several factors, e.g. the deformation of in-body cavities or the lack of texture. In this paper we present Endo-Depth-and-Motion, a pipeline that estimates the 6-degrees-of-freedom camera pose and dense

Cited by 161SourcecodeScholar
2021

Rotation-Only Bundle Adjustment

CVPR 2021poster

We propose a novel method for estimating the global rotations of the cameras independently of their positions and the scene structure. When two calibrated cameras observe five or more of the same points, their relative rotation can be recovered independently of the translation. We extend this idea t…

Cited by 20PDFScholar
2020

Corners for Layout: End-to-End Layout Recovery From 360 Images

RA-L 2020

The problem of 3D layout recovery in indoor scenes has been a core research topic for over a decade. However, there are still several major challenges that remain unsolved. Among the most relevant ones, a major part of the state-of-the-art methods make implicit or explicit assumptions on the scenes

Cited by 111SourceScholar
2020

From Points to Planes - Adding Planar Constraints to Monocular SLAM Factor Graphs

IROS 2020poster

Planar structures are common in man-made environments. Their addition to monocular SLAM algorithms is of relevance in order to achieve more complete and higher- level scene representations. Also, the additional constraints they introduce might reduce the estimation errors in certain situations. In t…

Cited by 27SourceScholar
2020

Mapillary Street-Level Sequences: A Dataset for Lifelong Place Recognition

CVPR 2020oral

Lifelong place recognition is an essential and challenging task in computer vision with vast applications in robust localization and efficient large-scale 3D reconstruction. Progress is currently hindered by a lack of large, diverse, publicly available datasets. We contribute with Mapillary Street-L…

Cited by 259PDFScholar
2019

CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth

CVPR 2019poster

Single-view depth estimation suffers from the problem that a network trained on images from one camera does not generalize to images taken with a different camera model. Thus, changing the camera model requires collecting an entirely new training dataset. In this work, we propose a new type of convo…

Cited by 173PDFScholar
2017

A multimodal dataset for object model learning from natural human-robot interaction

IROS 2017poster

Learning object models in the wild from natural human interactions is an essential ability for robots to perform general tasks. In this paper we present a robocentric multimodal dataset addressing this key challenge. Our dataset focuses on interactions where the user teaches new objects to the robot…

Cited by 20SourceScholar
2015

Layout aware visual tracking and mapping

IROS 2015poster

Nowadays real time visual Simultaneous Localization And Mapping (SLAM) algorithms exist and rely on consistent measurements across multiple views. In indoor environments, where majority of robot's activity takes place, severe occlusions can occur, e.g., when turning around a corner or moving from on…

Cited by 25SourceScholar
2015

Stereo parallel tracking and mapping for robot localization

IROS 2015poster

This paper describes a visual SLAM system based on stereo cameras and focused on real-time localization for mobile robots. To achieve this, it heavily exploits the parallel nature of the SLAM problem, separating the time-constrained pose estimation from less pressing matters such as map building and…

Cited by 175SourceScholar