← Search

Jakob Engel

14 accepted papers

2026

ART: Articulated Reconstruction Transformer

CVPR 2026

We introduce ART, Articulated Reconstruction Transformer--a category-agnostic, feed-forward model that reconstructs complete 3D articulated objects from only sparse, multi-state RGB images. Previous methods for articulated object reconstruction either rely on slow optimization with fragile cross-sta

Cited by 0SourceScholar
2026

JRM: Joint Reconstruction Model for Multiple Objects without Alignment

CVPR 2026

Object-centric reconstruction seeks to recover the 3D structure of a scene through composition of independent objects. While this independence can simplify modeling, it discards strong signals that could improve reconstruction, notably repetition where the same object model is seen multiple times in

Cited by 0SourceScholar
2026

LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World

CVPR 2026

Tracking 3D human motion from egocentric, multi-camera devices is challenged by severe egomotion and partial visibility or occlusions. Existing methods are designed for monocular video often recorded from static or slowly-moving cameras and cannot easily leverage multi-view, calibrated and localized

Cited by 0SourcecodeScholar
2026

ShapeR: Robust Conditional 3D Shape Generation from Casual Captures

CVPR 2026

Recent advances in 3D shape generation have achieved impressive results, but most existing methods rely on clean, unoccluded, and well-segmented inputs. Such conditions are rarely met in real-world scenarios. We present ShapeR, a novel approach for conditional 3D object shape generation from casuall

Cited by 0SourcecodeScholar
2025

Benchmarking Egocentric Visual-Inertial SLAM at City Scale

ICCV 2025poster

Precise 6-DoF simultaneous localization and mapping (SLAM) from onboard sensors is critical for wearable devices capturing egocentric data, which exhibits specific challenges, such as a wider diversity of motions and viewpoints, prevalent dynamic visual content, or long sessions affected by time-var…

Cited by 0SourcePDFScholar
2025

Human-in-the-Loop Local Corrections of 3D Scene Layouts via Infilling

ICCV 2025poster

We present a novel human-in-the-loop approach to estimate 3D scene layout that uses human feedback from an egocentric standpoint. We study this approach through introduction of a novel local correction task, where users identify local errors and prompt a model to automatically correct them. Building…

Cited by 0SourcePDFScholar
2025

Sonata: Self-Supervised Learning of Reliable Point Representations

CVPR 2025highlight

In this paper, we question whether we have a reliable self-supervised point cloud model that can be used for diverse 3D tasks via simple linear probing, even with limited data and minimal computation. We find that existing 3D self-supervised learning approaches fall short when evaluated on represent…

2025

VertexRegen: Mesh Generation with Continuous Level of Detail

ICCV 2025poster

We introduce VertexRegen, a novel mesh generation framework that enables generation at a continuous level of detail. Existing autoregressive methods generate meshes in a partial-to-complete manner and thus intermediate steps of generation represent incomplete structures. VertexRegen takes inspiratio…

Cited by 0SourcePDFScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Nymeria: A Massive Collection of Egocentric Multi-modal Human Motion in the Wild

ECCV 2024poster

"We introduce - a large-scale, diverse, richly annotated human motion dataset collected in the wild with multiple multimodal egocentric devices. The dataset comes with a) full-body ground-truth motion; b) multiple multimodal egocentric data from Project Aria devices with videos, eye tracking, IMUs a…

2024

SceneScript: Reconstructing Scenes With An Autoregressive Structured Language Model

ECCV 2024poster

"We introduce , a method that directly produces full scene models as a sequence of structured language commands using an autoregressive, token-based approach. Our proposed scene representation is inspired by recent successes in transformers & LLMs, and departs from more traditional methods which com…

Cited by 25SourcePDFScholar
2021

MR-iSAM2: Incremental Smoothing and Mapping with Multi-Root Bayes Tree for Multi-Robot SLAM

IROS 2021poster

We present multi-robot iSAM2 (MR-iSAM2), an efficient incremental smoothing and mapping (iSAM) algorithm to solve multi-robot simultaneous localization and mapping (SLAM) inference problems. MR-iSAM2 is based on a novel data structure multi-root Bayes tree (MRBT), which packs multiple Bayes trees wi…

Cited by 14SourceScholar