← Search

Martin R. Oswald

58 accepted papers

2026

Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

CVPR 2026

While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by introducing Chorus, a multi-teacher pretraining framework that learns a holistic feed-forward 3D Gaussian Splatting (3DG

Cited by 0SourcecodeScholar
2026

Edge-Centric Relational Reasoning for 3D Scene Graph Prediction

AAAI 2026technical

3D scene graph prediction aims to abstract complex 3D environments into structured graphs consisting of objects and their pairwise relationships. Existing approaches typically adopt object-centric graph neural networks, where relation edge features are iteratively updated by aggregating messages fro

Cited by 0SourcePDFScholar
2026

From Rays to Projections: Better Inputs for Feed-Forward View Synthesis

CVPR 2026

Feed-forward view synthesis models predict a novel view in a single pass with minimal 3D inductive bias. Existing works encode cameras as Plucker ray maps, which tie predictions to the arbitrary world coordinate gauge and make them sensitive to small camera transformations, thereby undermining geome

Cited by 0SourcecodeScholar
2026

Gaussian Mapping for Evolving Scenes

CVPR 2026

Mapping systems with novel view synthesis (NVS) capabilities are widely used in computer vision, as well as in various applications, including augmented reality, robotics, and autonomous driving. Most notably, 3D Gaussian Splatting-based systems show high NVS performance; however, many current appro

Cited by 0SourcecodeScholar
2026

MCGS-SLAM: A Multi-Camera SLAM Framework Using Gaussian Splatting for High-Fidelity Mapping

ICRA 2026poster

Recent progress in dense SLAM has primarily targeted monocular setups, often at the expense of robustness and geometric coverage. We present MCGS-SLAM, the first purely RGB-based multi-camera SLAM system built on 3D Gaussian Splatting (3DGS). Unlike prior methods relying on sparse maps or inertial d…

2026

OrthoRF: Exploring Orthogonality in Object-Centric Representations

ICLR 2026poster

Neural synchrony is hypothesized to help the brain organize visual scenes into structured multi-object representations. In machine learning, synchrony-based models analogously learn object-centric representations by storing binding in the phase of complex-valued features. Rotating Features (RF) inst…

Cited by 0SourceScholar
2026

Unblur-SLAM: Dense Neural SLAM for Blurry Inputs

CVPR 2026

We propose Unblur-SLAM, an RGB SLAM pipeline for sharp 3D reconstruction from blurred image inputs. In contrast to previous work, our approach is able to handle different types of blur and demonstrates state-of-the-art performance in the presence of both motion blur and defocus blur. Moreover, we ad

Cited by 0SourcecodeScholar
2025

3D-AVS: LiDAR-based 3D Auto-Vocabulary Segmentation

CVPR 2025poster

Open-vocabulary segmentation methods offer promising capabilities in detecting unseen object categories, but the category must be aware and needs to be provided by a human, either via a text prompt or pre-labeled datasets, thus limiting their scalability. We propose 3D-AVS, a method for Auto-Vocabul…

2025

An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels

ICLR 2025poster

This work does not introduce a new method. Instead, we present an interesting finding that questions the necessity of the inductive bias of locality in modern computer vision architectures. Concretely, we find that vanilla Transformers can operate by directly treating each individual pixel as a toke…

Cited by 13SourcePDFScholar
2025

ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular Videos

CVPR 2025poster

Creating a photorealistic scene and human reconstruction from a single monocular in-the-wild video figures prominently in the perception of a human-centric 3D world. Recent neural rendering advances have enabled holistic human-scene reconstruction but require pre-calibrated camera and human poses, a…

2025

ProDyG: Progressive Dynamic Scene Reconstruction via Gaussian Splatting from Monocular Videos

NeurIPS 2025poster

Achieving truly practical dynamic 3D reconstruction requires online operation, global pose and map consistency, detailed appearance modeling, and the flexibility to handle both RGB and RGB-D inputs. However, existing SLAM methods typically merely remove the dynamic parts or require RGB-D input, whil…

Cited by 0SourceScholar
2025

SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting

NeurIPS 2025poster

3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven to be an effective strategy for 3D scene understanding. Current Language Gaussian Splatting line of work fall into thre…

Cited by 0SourceScholar
2025

SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

ICCV 2025poster

Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training, or together at inference. This highlights a clear absence of a model capable of processing 3D data…

2025

Splat-LOAM: Gaussian Splatting LiDAR Odometry and Mapping

ICCV 2025poster

LiDARs provide accurate geometric measurements, making them valuable for ego-motion estimation and reconstruction tasks.Although its success, managing an accurate and lightweight representation of the environment still poses challenges.Both classic and NeRF-based solutions have to trade off accuracy…

2025

TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning

ICCV 2025poster

Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Multimodal Large Language Models (MLLMs) struggle at this task. In this paper, we introduce TWIST & SCOUT, a framework that equips pre-trained MLLMs with visual grounding abil…

Cited by 0SourcePDFScholar
2025

ToF-Splatting: Dense SLAM using Sparse Time-of-Flight Depth and Multi-Frame Integration

ICCV 2025poster

Time-of-Flight (ToF) sensors provide efficient active depth sensing at relatively low power budgets; among such designs, only very sparse measurements from low-resolution sensors are considered to meet the increasingly limited power constraints of mobile and AR/VR devices. However, such extreme spar…

Cited by 0SourcePDFScholar
2025

Union-over-Intersections: Object Detection beyond Winner-Takes-All

ICLR 2025spotlight

This paper revisits the problem of predicting box locations in object detection architectures. Typically, each box proposal or box query aims to directly maximize the intersection-over-union score with the ground truth, followed by a winner-takes-all non-maximum suppression where only the highest sc…

2024

Learning High-level Semantic-Relational Concepts for SLAM

IROS 2024poster

Recent works on SLAM extend their pose graphs with higher-level semantic concepts like Rooms exploiting relationships between them, to provide, not only a richer representation of the situation/environment but also to improve the accuracy of its estimation. Concretely, our previous work, Situational…

Cited by 1SourceScholar
2024

Loopy-SLAM: Dense Neural SLAM with Loop Closures

CVPR 2024poster

Neural RGBD SLAM techniques have shown promise in dense Simultaneous Localization And Mapping (SLAM) yet face challenges such as error accumulation during camera tracking resulting in distorted maps. In response we introduce Loopy-SLAM that globally optimizes poses and the dense 3D model. We use fra…

2024

R-MAE: Regions Meet Masked Autoencoders

ICLR 2024poster

In this work, we explore regions as a potential visual analogue of words for self-supervised image representation learning. Inspired by Masked Autoencoding (MAE), a generative pre-training baseline, we propose masked region autoencoding to learn from groups of pixels or regions. Specifically, we des…

2024

WorldPose: A World Cup Dataset for Global 3D Human Pose Estimation

ECCV 2024poster

"We present , a novel dataset for advancing research in multi-person global pose estimation in the wild, featuring footage from the 2022 FIFA World Cup. While previous datasets have primarily focused on local poses, often limited to a single person or in constrained, indoor settings, the infrastruct…

Cited by 5SourcePDFScholar
2023

DeepLSD: Line Segment Detection and Refinement With Deep Image Gradients

CVPR 2023poster

Line segments are ubiquitous in our human-made world and are increasingly used in vision tasks. They are complementary to feature points thanks to their spatial extent and the structural information they provide. Traditional line detectors based on the image gradient are extremely fast and accurate,…

2023

Detecting Objects with Context-Likelihood Graphs and Graph Refinement

ICCV 2023poster

The goal of this paper is to detect objects by exploiting their interrelationships. Contrary to existing methods, which learn objects and relations separately, our key idea is to learn the object-relation distribution jointly. We first propose a novel way of creating a graphical representation of an…

Cited by 2PDFScholar
2023

Human from Blur: Human Pose Tracking from Blurry Images

ICCV 2023poster

We propose a method to estimate 3D human poses from substantially blurred images. The key idea is to tackle the inverse problem of image deblurring by modeling the forward problem with a 3D human model, a texture map, and a sequence of poses to describe human motion. The blurring process is then mod…

Cited by 3PDFScholar
2023

Learning-based Relational Object Matching Across Views

ICRA 2023poster

Intelligent robots require object-level scene understanding to reason about possible tasks and interactions with the environment. Moreover, many perception tasks such as scene reconstruction, image retrieval, or place recognition can benefit from reasoning on the level of objects. While keypoint-bas…

Cited by 5SourceScholar
2023

The Drunkard’s Odometry: Estimating Camera Motion in Deforming Scenes

NeurIPS 2023poster

Estimating camera motion in deformable scenes poses a complex and open research challenge. Most existing non-rigid structure from motion techniques assume to observe also static scene parts besides deforming scene parts in order to establish an anchoring reference. However, this assumption does not…

2023

Tracking by 3D Model Estimation of Unknown Objects in Videos

ICCV 2023poster

Most model-free visual object tracking methods formulate the tracking task as object location estimation given by a 2D segmentation or a bounding box in each video frame. We argue that this representation is limited and instead propose to guide and improve 2D tracking with an explicit object represe…

Cited by 7PDFScholar
2022

A Real-Time Online Learning Framework for Joint 3D Reconstruction and Semantic Segmentation of Indoor Scenes

RA-L 2022

This letter presents a real-time online vision framework to jointly recover an indoor scene’s 3D structure and semantic label. Given noisy depth maps, a camera trajectory, and 2D semantic labels at train time, the proposed deep neural network based approach learns to fuse the depth over frames with

Cited by 26SourcecodeScholar
2022

BoxeR: Box-Attention for 2D and 3D Transformers

CVPR 2022poster

In this paper, we propose a simple attention mechanism, we call Box-Attention. It enables spatial interaction between grid features, as sampled from boxes of interest, and improves the learning capability of transformers for several vision tasks. Specifically, we present BoxeR, short for Box Transfo…

Cited by 43PDFcodeScholar
2022

CompNVS: Novel View Synthesis with Scene Completion

ECCV 2022poster

"We introduce a scalable framework for novel view synthesis from RGB-D images with largely incomplete scene coverage. While generative neural approaches have demonstrated spectacular results on 2D images, they have not yet achieved similar photorealistic results in combination with scene completion…

Cited by 8SourcePDFScholar
2022

Learning Online Multi-sensor Depth Fusion

ECCV 2022poster

"Many hand-held or mixed reality devices are used with a single sensor for 3D reconstruction, although they often comprise multiple sensors. Multi-sensor depth fusion is able to substantially improve the robustness and accuracy of 3D reconstruction methods, but existing techniques are not robust eno…

2022

Motion-From-Blur: 3D Shape and Motion Estimation of Motion-Blurred Objects in Videos

CVPR 2022poster

We propose a method for jointly estimating the 3D motion, 3D shape, and appearance of highly motion-blurred objects from a video. To this end, we model the blurred appearance of a fast moving object in a generative fashion by parametrizing its 3D position, rotation, velocity, acceleration, bounces,…

Cited by 10PDFcodeScholar
2022

NICE-SLAM: Neural Implicit Scalable Encoding for SLAM

CVPR 2022poster

Neural implicit representations have recently shown encouraging results in various domains, including promising progress in simultaneous localization and mapping (SLAM). Nevertheless, existing methods produce over-smoothed scene reconstructions and have difficulty scaling up to large scenes. These l…

Cited by 766PDFcodeScholar
2021

DeFMO: Deblurring and Shape Recovery of Fast Moving Objects

CVPR 2021poster

Objects moving at high speed appear significantly blurred when captured with cameras. The blurry appearance is especially ambiguous when the object has complex shape or texture. In such cases, classical methods, or even humans, are unable to recover the object's appearance and motion. We propose a m…

Cited by 50PDFcodeScholar
2021

FMODetect: Robust Detection of Fast Moving Objects

ICCV 2021poster

We propose the first learning-based approach for fast moving objects detection. Such objects are highly blurred and move over large distances within one video frame. Fast moving objects are associated with a deblurring and matting problem, also called deblatting. We show that the separation of debla…

Cited by 14PDFcodeScholar
2021

NeuralFusion: Online Depth Fusion in Latent Space

CVPR 2021poster

We present a novel online depth map fusion approach that learns depth map aggregation in a latent feature space. While previous fusion methods use an explicit scene representation like signed distance functions (SDFs), we propose a learned feature representation for the fusion. The key idea is a sep…

Cited by 65PDFcodeScholar
2021

SOLD2: Self-Supervised Occlusion-Aware Line Description and Detection

CVPR 2021poster

Compared to feature point detection and description, detecting and matching line segments offer additional challenges. Yet, line features represent a promising complement to points for multi-view tasks. Lines are indeed well-defined by the image gradient, frequently appear even in poorly textured ar…

Cited by 97PDFcodeScholar
2021

Sat2Vid: Street-View Panoramic Video Synthesis From a Single Satellite Image

ICCV 2021poster

We present a novel method for synthesizing both temporally and geometrically consistent street-view panoramic video from a single satellite image and camera trajectory. Existing cross-view synthesis approaches focus on images, while video synthesis in such a case has not yet received enough attentio…

Cited by 12PDFScholar
2021

Shape from Blur: Recovering Textured 3D Shape and Motion of Fast Moving Objects

NeurIPS 2021poster

We address the novel task of jointly reconstructing the 3D shape, texture, and motion of an object from a single motion-blurred image. While previous approaches address the deblurring problem only in the 2D image domain, our proposed rigorous modeling of all object properties in the 3D domain enable…

2020

Accurate Mapping and Planning for Autonomous Racing

IROS 2020poster

This paper presents the perception, mapping, and planning pipeline implemented on an autonomous race car. It was developed by the 2019 AMZ driverless team for the Formula Student Germany (FSG) 2019 driverless competition, where it won 1st place overall. The presented solution combines early fusion o…

Cited by 31SourceScholar
2020

Aerial Single-View Depth Completion With Image-Guided Uncertainty Estimation

RA-L 2020

On the pursuit of autonomous flying robots, the scientific community has been developing onboard real-time algorithms for localisation, mapping and planning. Despite recent progress, the available solutions still lack accuracy and robustness in many aspects. While mapping for autonomous cars had a s

Cited by 60SourcecodeScholar
2020

Geometry-Aware Satellite-to-Ground Image Synthesis for Urban Areas

CVPR 2020poster

We present a novel method for generating panoramic street-view images which are geometrically consistent with a given satellite image. Different from existing approaches that completely rely on a deep learning architecture to generalize cross-view image distributions, our approach explicitly loops i…

Cited by 78PDFScholar
2020

Online Invariance Selection for Local Feature Descriptors

ECCV 2020poster

To be invariant, or not to be invariant: that is the question formulated in this work about local descriptors. A limitation of current feature descriptors is the trade-off between generalization and discriminative power: more invariance means less informative descriptors. We propose to overcome this…

2020

RoutedFusion: Learning Real-Time Depth Map Fusion

CVPR 2020oral

The efficient fusion of depth maps is a key part of most state-of-the-art 3D reconstruction methods. Besides requiring high accuracy, these depth fusion methods need to be scalable and real-time capable. To this end, we present a novel real-time capable machine learning-based method for depth map fu…

Cited by 97PDFcodeScholar
2018

Consensus Maximization for Semantic Region Correspondences

CVPR 2018poster

We propose a novel method for the geometric registration of semantically labeled regions. We approximate semantic regions by ellipsoids, and leverage their convexity to formulate the correspondence search effectively as a constrained optimization problem that maximizes the number of matched regions,…

Cited by 8SourcePDFScholar
2018

Learning Priors for Semantic 3D Reconstruction

ECCV 2018poster

We present a novel semantic 3D reconstruction framework which embeds variational regularization into a neural network. Our network performs a fixed number of unrolled multi-scale optimization iterations with shared interaction weights. In contrast to existing variational methods for semantic 3D reco…

Cited by 56SourcePDFScholar
2017

Consensus Maximization With Linear Matrix Inequality Constraints

CVPR 2017poster

Consensus maximization has proven to be a useful tool for robust estimation. While randomized methods like RANSAC are fast, they do not guarantee global optimality and fail to manage large amounts of outliers. On the other hand, global methods are commonly slow because they do not exploit the struct…

Cited by 42PDFScholar
2017

Indoor Scan2BIM: Building information models of house interiors

IROS 2017poster

We present a system to generate building information models (BIMs) of house interiors from 3D scans. The strength of our approach is its simplicity and low runtime which allows for mobile processing applications. We consider scans of single floor, Manhattan-like indoor scenes for which our method cr…

Cited by 87SourceScholar
2017

Semantically Informed Multiview Surface Refinement

ICCV 2017poster

We present a method to jointly refine the geometry and semantic segmentation of 3D surface meshes. Our method alternates between updating the shape and the semantic labels. In the geometry refinement step, the mesh is deformed with variational energy minimization, such that it simultaneously maximiz…

Cited by 37PDFScholar
2015

Entropy Minimization for Convex Relaxation Approaches

ICCV 2015poster

Despite their enormous success in solving hard combinatorial problems, convex relaxation approaches often suffer from the fact that the computed solutions are far from binary and that subsequent heuristic binarization may substantially degrade the quality of computed solutions. In this paper, we pr…

Cited by 7PDFScholar