← Search

Ondrej Miksik

15 accepted papers

2026

UnLoc: Leveraging Depth Uncertainties for Floorplan Localization

ICLR 2026poster

We propose UnLoc, an efficient data-driven solution for sequential camera localization within floorplans. Floorplan data is readily available, long-term persistent, and robust to changes in visual appearance. We address key limitations of recent methods, such as the lack of uncertainty modeling in d…

Cited by 0SourcecodeScholar
2025

CrossOver: 3D Scene Cross-Modal Alignment

CVPR 2025highlight

Multi-modal 3D object understanding has gained significant attention, yet current approaches often assume complete data availability and rigid alignment across all modalities. We present CrossOver, a novel framework for cross-modal 3D scene understanding via flexible, scene-level modality alignment.…

2023

SGAligner: 3D Scene Alignment with Scene Graphs

ICCV 2023poster

Building 3D scene graphs has recently emerged as a topic in scene representation for several embodied AI applications to represent the world in a structured and rich manner. With their increased use in solving downstream tasks (e.g., navigation and room rearrangement), can we leverage and recycle th…

Cited by 14PDFcodeScholar
2022

LaMAR: Benchmarking Localization and Mapping for Augmented Reality

ECCV 2022poster

"Localization and mapping is the foundational technology for augmented reality (AR) that enables sharing and persistence of digital content in the real world. While significant progress has been made, researchers are still mostly driven by unrealistic benchmarks not representative of real-world AR s…

2022

Learning To Detect Scene Landmarks for Camera Localization

CVPR 2022oral

Modern camera localization methods that use image retrieval, feature matching, and 3D structure-based pose estimation require long-term storage of numerous scene images or a vast amount of image features. This can make them unsuitable for resource constrained VR/AR devices and also raises serious pr…

Cited by 40PDFcodeScholar
2022

Learning to Simulate Realistic LiDARs

IROS 2022poster

Simulating realistic sensors is a challenging part in data generation for autonomous systems, often involving carefully handcrafted sensor design, scene properties, and physics modeling. To alleviate this, we introduce a pipeline for data-driven simulation of a realistic LiDAR sensor. We propose a m…

Cited by 19SourceScholar
2021

Cross-Descriptor Visual Localization and Mapping

ICCV 2021poster

Visual localization and mapping is the key technology underlying the majority of mixed reality and robotics systems. Most state-of-the-art approaches rely on local features to establish correspondences between images. In this paper, we present three novel scenarios for localization and mapping which…

Cited by 35PDFcodeScholar
2021

Taskography: Evaluating robot task planning over large 3D scene graphs

CoRL 2021poster

3D scene graphs (3DSGs) are an emerging description; unifying symbolic, topological, and metric scene representations. However, typical 3DSGs contain hundreds of objects and symbols even for small environments; rendering task planning on the \emph{full} graph impractical. We construct \textbf{Taskog…

Cited by 84SourcecodeScholar
2018

On the Robustness of Semantic Segmentation Models to Adversarial Attacks

CVPR 2018poster

Deep Neural Networks (DNNs) have been demonstrated to perform exceptionally well on most recognition tasks such as image classification and segmentation. However, they have also been shown to be vulnerable to adversarial examples. This phenomenon has recently attracted a lot of attention but it has…

2018

Real-Time Dense Stereo Matching With ELAS on FPGA-Accelerated Embedded Devices

RA-L 2018

For many applications in low-power real-time robotics, stereo cameras are the sensors of choice for depth perception as they are typically cheaper and more versatile than their active counterparts. Their biggest drawback, however, is that they do not directly sense depth maps; instead, these must be

Cited by 33SourcecodeScholar
2017

ROAM: A Rich Object Appearance Model With Application to Rotoscoping

CVPR 2017poster

Rotoscoping, the detailed delineation of scene elements through a video shot, is a painstaking task of tremendous importance in professional post-production pipelines. While pixel-wise segmentation techniques can help for this task, professional rotoscoping tools rely on parametric curves that offer…

Cited by 6PDFScholar
2016

Staple: Complementary Learners for Real-Time Tracking

CVPR 2016poster

Correlation Filter-based trackers have recently achieved excellent performance, showing great robustness to challenging situations exhibiting motion blur and illumination changes. However, since the model that they learn depends strongly on the spatial layout of the tracked object, they are notoriou…

Cited by 2208PDFScholar
2015

Incremental dense multi-modal 3D scene reconstruction

IROS 2015poster

Aquiring reliable depth maps is an essential prerequisite for accurate and incremental 3D reconstruction used in a variety of robotics applications. Depth maps produced by affordable Kinect-like cameras have become a de-facto standard for indoor reconstruction and the driving force behind the succes…

Cited by 19SourceScholar
2015

Incremental dense semantic stereo fusion for large-scale semantic scene reconstruction

ICRA 2015poster

Our abilities in scene understanding, which allow us to perceive the 3D structure of our surroundings and intuitively recognise the objects we see, are things that we largely take for granted, but for robots, the task of understanding large scenes quickly remains extremely challenging. Recently, sce…

Cited by 260SourceScholar