← Search

Francis Engelmann

31 accepted papers

2026

FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

CVPR 2026

We present FunREC, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunREC operates directly on in-t

Cited by 0SourcecodeScholar
2026

HoMeR: Learning In-The-Wild Mobile Manipulation Via Hybrid Imitation and Whole-Body Control

ICRA 2026poster

We introduce HoMeR, an imitation learning framework for mobile manipulation that combines whole-body control with hybrid action modes that handle both long-range and fine-grained motion, enabling effective performance on realistic in-the-wild tasks. At its core is a fast, kinematics-based whole-body…

2026

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

ICRA 2026poster

Open-vocabulary 3D segmentation enables the exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying object-level instances in a scene. However, they face challenges when it comes to understanding more fi…

2026

SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling

ICLR 2026poster

Generative methods for 3D assets have recently achieved remarkable progress, yet providing intuitive and precise control over the object geometry remains a key challenge. Existing approaches predominantly rely on text or image prompts, which often fall short in geometric specificity: language can be…

Cited by 0SourceScholar
2026

UnLoc: Leveraging Depth Uncertainties for Floorplan Localization

ICLR 2026poster

We propose UnLoc, an efficient data-driven solution for sequential camera localization within floorplans. Floorplan data is readily available, long-term persistent, and robust to changes in visual appearance. We address key limitations of recent methods, such as the lack of uncertainty modeling in d…

Cited by 0SourcecodeScholar
2025

ARKit LabelMaker: A New Scale for Indoor 3D Scene Understanding

CVPR 2025poster

Neural network performance scales with both model size and data volume, as shown in both language and image processing. This requires scaling-friendly architectures and large datasets. While transformers have been adapted for 3D vision, a `GPT-moment' remains elusive due to limited training data. We…

2025

HouseLayout3D: A Benchmark and Training-free Baseline for 3D Layout Estimation in the Wild

NeurIPS 2025poster

Current 3D layout estimation models are predominantly trained on synthetic datasets biased toward simplistic, single-floor scenes. This prevents them from generalizing to complex, multi-floor buildings, often forcing a per-floor processing approach that sacrifices global context. Few works have atte…

Cited by 0SourceScholar
2025

LookOut: Real-World Humanoid Egocentric Navigation

ICCV 2025poster

The ability to predict collision-free future trajectories from egocentric observations is crucial in applications such as humanoid robotics, VR / AR, and assistive navigation. In this work, we introduce the challenging problem of predicting a sequence of future 6D head poses from an egocentric video…

2025

Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces

CVPR 2025highlight

We introduce the task of predicting functional 3D scene graphs for real-world indoor environments from posed RGB-D images. Unlike traditional 3D scene graphs that focus on spatial relationships of objects, functional 3D scene graphs capture objects, interactive elements, and their functional relatio…

2025

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

RA-L 2025

Open-vocabulary 3D segmentation enables exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">object</i

Cited by 31SourceScholar
2025

SuperDec: 3D Scene Decomposition with Superquadrics Primitives

ICCV 2025poster

We present SuperDec, an approach for compact 3D scene representations based on geometric primitives, namely superquadrics.While most recent works leverage geometric primitives to obtain photorealistic 3D scene representations, we propose to leverage them to obtain a compact yet expressive representa…

Cited by 0SourcePDFScholar
2025

Video Perception Models for 3D Scene Synthesis

NeurIPS 2025poster

Automating the expert-dependent and labor-intensive task of 3D scene synthesis would significantly benefit fields such as architectural design, robotics simulation, and virtual reality. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or…

Cited by 0SourceScholar
2024

AGILE3D: Attention Guided Interactive Multi-object 3D Segmentation

ICLR 2024poster

During interactive segmentation, a model and a user work together to delineate objects of interest in a 3D point cloud. In an iterative process, the model assigns each data point to an object (or the background), while the user corrects errors in the resulting segmentation and feeds them back into t…

2024

ICGNet: A Unified Approach for Instance-Centric Grasping

ICRA 2024poster

Accurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of scene understanding: First, the robot needs to analyze the geometric properties of individual objects to find feasible…

Cited by 13SourcecodeScholar
2024

Improving 2D Feature Representations by 3D-Aware Fine-Tuning

ECCV 2024poster

"Current visual foundation models are trained purely on unstructured 2D data, limiting their understanding of 3D structure of objects and scenes. In this work, we show that fine-tuning on 3D-aware data improves the quality of emerging semantic features. We design a method to lift semantic 2D feature…

2024

OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views

ICLR 2024poster

Large visual-language models (VLMs), like CLIP, enable open-set image segmentation to segment arbitrary concepts from an image in a zero-shot manner. This goes beyond the traditional closed-set assumption, i.e., where models can only segment classes from a pre-defined training set. More recently, fi…

Cited by 33SourcePDFScholar
2024

SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D Scenes

CVPR 2024poster

Existing 3D scene understanding methods are heavily focused on 3D semantic and instance segmentation. However identifying objects and their parts only constitutes an intermediate step towards a more fine-grained goal which is effectively interacting with the functional interactive elements (e.g. han…

Cited by 35SourcePDFScholar
2024

SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs

ECCV 2024poster

"We introduce the task of localizing an input image within a multi-modal reference map represented by a collection of 3D scene graphs. These scene graphs comprise multiple modalities, including object-level point clouds, images, attributes, and relationships between objects, offering a lightweight a…

2024

Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels

ECCV 2024poster

"Current 3D scene segmentation methods are heavily dependent on manually annotated 3D training datasets. Such manual annotations are labor-intensive, and often lack fine-grained details. Furthermore, models trained on this data typically struggle to recognize object classes beyond the annotated trai…

Cited by 34SourcePDFScholar
2023

3D Segmentation of Humans in Point Clouds with Synthetic Data

ICCV 2023poster

Segmenting humans in 3D indoor scenes has become increasingly important with the rise of human-centered robotics and AR/VR applications. To this end, we propose the task of joint 3D human semantic segmentation, instance segmentation and multi-human body-part segmentation. Few works have attempted to…

Cited by 29PDFScholar
2023

Connecting the Dots: Floorplan Reconstruction Using Two-Level Queries

CVPR 2023poster

We address 2D floorplan reconstruction from 3D scans. Existing approaches typically employ heuristically designed multi-stage pipelines. Instead, we formulate floorplan reconstruction as a single-stage structured prediction task: find a variable-size set of polygons, which in turn are variable-lengt…

2023

Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

ICRA 2023poster

Modern 3D semantic instance segmentation approaches predominantly rely on specialized voting mechanisms followed by carefully designed geometric clustering techniques. Building on the successes of recent Transformer-based methods for object detection and image segmentation, we propose the first Tran…

Cited by 272SourcecodeScholar
2023

OpenMask3D: Open-Vocabulary 3D Instance Segmentation

NeurIPS 2023poster

We introduce the task of open-vocabulary 3D instance segmentation. Current approaches for 3D instance segmentation can typically only recognize object categories from a pre-defined closed set of classes that are annotated in the training datasets. This results in important limitations for real-world…

2022

Box2Mask: Weakly Supervised 3D Semantic Instance Segmentation Using Bounding Boxes

ECCV 2022poster

"Current 3D segmentation methods heavily rely on large-scale point-cloud datasets, which are notoriously laborious to annotate. Few attempts have been made to circumvent the need for dense per-point annotations. In this work, we look at weakly-supervised 3D semantic instance segmentation. The key id…

Cited by 76SourcePDFScholar
2020

3D-MPA: Multi-Proposal Aggregation for 3D Semantic Instance Segmentation

CVPR 2020poster

We present 3D-MPA, a method for instance segmentation on 3D point clouds. Given an input point cloud, we propose an object-centric approach where each point votes for its object center. We sample object proposals from the predicted object centers. Then, we learn proposal features from grouped point…

Cited by 251PDFScholar
2020

Dilated Point Convolutions: On the Receptive Field Size of Point Convolutions on 3D Point Clouds

ICRA 2020poster

In this work, we propose Dilated Point Convolutions (DPC). In a thorough ablation study, we show that the receptive field size is directly related to the performance of 3D point cloud processing tasks, including semantic segmentation and object classification. Point convolutions are widely used to e…

Cited by 114SourceScholar
2020

DualConvMesh-Net: Joint Geodesic and Euclidean Convolutions on 3D Meshes

CVPR 2020oral

We propose DualConvMesh-Nets (DCM-Net) a family of deep hierarchical convolutional networks over 3D geometric data that combines two types of convolutions. The first type, Geodesic convolutions, defines the kernel weights over mesh surfaces or graphs. That is, the convolutional kernel weights are ma…

Cited by 107PDFcodeScholar
2017

Keyframe-based visual-inertial online SLAM with relocalization

IROS 2017poster

Complementing images with inertial measurements has become one of the most popular approaches to achieve highly accurate and robust real-time camera pose tracking. In this paper, we present a keyframe-based approach to visual-inertial simultaneous localization and mapping (SLAM) for monocular and st…

Cited by 63SourceScholar
2016

Multi-scale object candidates for generic object tracking in street scenes

ICRA 2016

Most vision based systems for object tracking in urban environments focus on a limited number of important object categories such as cars or pedestrians, for which powerful detectors are available. However, practical driving scenarios contain many additional objects of interest, for which suitable d

Cited by 44SourceScholar