← Search

Johanna Wald

11 accepted papers

2026

OVI-MAP: Open-Vocabulary Instance-Semantic Mapping

CVPR 2026

Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, real-time processing, and flexible open-set reasoning. Existing methods often rely

Cited by 0SourcecodeScholar
2026

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

ICRA 2026poster

Open-vocabulary 3D segmentation enables the exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying object-level instances in a scene. However, they face challenges when it comes to understanding more fi…

2025

RelationField: Relate Anything in Radiance Fields

CVPR 2025poster

Neural radiance fields are an emerging 3D scene representation and recently even been extended to learn features for scene understanding by distilling open-vocabulary features from vision-language models. However, current method primarily focus on object-centric representations, supporting object se…

2025

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

RA-L 2025

Open-vocabulary 3D segmentation enables exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">object</i

Cited by 31SourceScholar
2023

Towards Long-Term Retrieval-Based Visual Localization in Indoor Environments With Changes

RA-L 2023

Visual localization is a challenging task due to the presence of illumination changes, occlusion, and perception from novel viewpoints. Re-localizing the camera pose in long-term setups raises difficulties caused by changes in scene appearance and geometry introduced by human or natural deterioratio

Cited by 12SourceScholar
2021

SceneGraphFusion: Incremental 3D Scene Graph Prediction From RGB-D Sequences

CVPR 2021poster

Scene graphs are a compact and explicit representation successfully used in a variety of 2D scene understanding tasks. This work proposes a method to build up semantic scene graphs from a 3D environment incrementally given a sequence of RGB-D frames. To this end, we aggregate PointNet features from…

Cited by 186PDFScholar
2020

Beyond Controlled Environments: 3D Camera Re-Localization in Changing Indoor Scenes

ECCV 2020poster

Long-term camera re-localization is an important task with numerous computer vision and robotics applications. Whilst various outdoor benchmarks exist that target lighting, weather and seasonal changes, far less attention has been paid to appearance changes that occur indoors. This has led to a mism…

2020

Learning 3D Semantic Scene Graphs From 3D Indoor Reconstructions

CVPR 2020poster

Scene understanding has been of high interest in computer vision. It encompasses not only identifying objects in a scene, but also their relationships within the given context. With this goal, a recent line of works tackles 3D semantic segmentation and scene layout prediction. In our work we focus o…

Cited by 262PDFScholar
2019

RIO: 3D Object Instance Re-Localization in Changing Indoor Environments

ICCV 2019oral

In this work, we introduce the task of 3D object instance re-localization (RIO): given one or multiple objects in an RGB-D scan, we want to estimate their corresponding 6DoF poses in another 3D scan of the same environment taken at a later point in time. We consider RIO a particularly important task…

Cited by 171PDFcodeScholar
2018

Fully-Convolutional Point Networks for Large-Scale Point Clouds

ECCV 2018poster

This work proposes a general-purpose, fully-convolutional network architecture for efficiently processing large-scale 3D data. One striking characteristic of our approach is its ability to process unorganized 3D representations such as point clouds as input, then transforming them internally to orde…

2018

Real-Time Fully Incremental Scene Understanding on Mobile Platforms

RA-L 2018

We propose an online RGB-D based scene understanding method for indoor scenes running in real time on mobile devices. First, we incrementally reconstruct the scene via simultaneous localization and mapping and compute a three-dimensional (3-D) geometric segmentation by fusing segments obtained from

Cited by 30SourceScholar