← Search

Richard Newcombe

32 accepted papers

2026

ART: Articulated Reconstruction Transformer

CVPR 2026

We introduce ART, Articulated Reconstruction Transformer--a category-agnostic, feed-forward model that reconstructs complete 3D articulated objects from only sparse, multi-state RGB images. Previous methods for articulated object reconstruction either rely on slow optimization with fragile cross-sta

Cited by 0SourceScholar
2026

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

RSS 2026poster

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limi…

Cited by 0SourceScholar
2026

JRM: Joint Reconstruction Model for Multiple Objects without Alignment

CVPR 2026

Object-centric reconstruction seeks to recover the 3D structure of a scene through composition of independent objects. While this independence can simplify modeling, it discards strong signals that could improve reconstruction, notably repetition where the same object model is seen multiple times in

Cited by 0SourceScholar
2026

LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World

CVPR 2026

Tracking 3D human motion from egocentric, multi-camera devices is challenged by severe egomotion and partial visibility or occlusions. Existing methods are designed for monocular video often recorded from static or slowly-moving cameras and cannot easily leverage multi-view, calibrated and localized

Cited by 0SourcecodeScholar
2026

ShapeR: Robust Conditional 3D Shape Generation from Casual Captures

CVPR 2026

Recent advances in 3D shape generation have achieved impressive results, but most existing methods rely on clean, unoccluded, and well-segmented inputs. Such conditions are rarely met in real-world scenarios. We present ShapeR, a novel approach for conditional 3D object shape generation from casuall

Cited by 0SourcecodeScholar
2025

4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos

NeurIPS 2025spotlight

We propose 4DGT, a 4D Gaussian-based Transformer model for dynamic scene reconstruction, trained entirely on real-world monocular posed videos. Using 4D Gaussian as an inductive bias, 4DGT unifies static and dynamic components, enabling the modeling of complex, time-varying environments with varying…

Cited by 0SourceScholar
2025

Benchmarking Egocentric Visual-Inertial SLAM at City Scale

ICCV 2025poster

Precise 6-DoF simultaneous localization and mapping (SLAM) from onboard sensors is critical for wearable devices capturing egocentric data, which exhibits specific challenges, such as a wider diversity of motions and viewpoints, prevalent dynamic visual content, or long sessions affected by time-var…

Cited by 0SourcePDFScholar
2025

DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

NeurIPS 2025poster

We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward scene reconstruction has gained significant attention for its ability to rapidly cr…

Cited by 0SourceScholar
2025

Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset

CVPR 2025highlight

We introduce Digital Twin Catalog (DTC), a new large-scale photorealistic 3D object digital twin dataset. A digital twin of a 3D object is a highly detailed, virtually indistinguishable representation of a physical object, accurately capturing its shape, appearance, physical properties, and other at…

2025

EgoLM: Multi-Modal Language Model of Egocentric Motions

CVPR 2025poster

As wearable devices become more prevalent, understanding the user's motion is crucial for improving contextual AI systems. We introduce EgoLM, a versatile framework designed for egocentric motion understanding using multi-modal data. EgoLM integrates the rich contextual information from egocentric v…

Cited by 4SourcePDFScholar
2025

HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos

CVPR 2025highlight

We introduce HOT3D, a publicly available dataset for egocentric hand and object tracking in 3D. The dataset offers over 833 minutes (3.7M+ images) of recordings that feature 19 subjects interacting with 33 diverse rigid objects. In addition to simple pick-up, observe, and put-down actions, the subje…

2025

Human-in-the-Loop Local Corrections of 3D Scene Layouts via Infilling

ICCV 2025poster

We present a novel human-in-the-loop approach to estimate 3D scene layout that uses human feedback from an egocentric standpoint. We study this approach through introduction of a novel local correction task, where users identify local errors and prompt a model to automatically correct them. Building…

Cited by 0SourcePDFScholar
2025

LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields

CVPR 2025poster

We present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent Large Reconstruction Models (LRMs) that achieve state-of-the-ar…

Cited by 0SourcePDFScholar
2025

Reading Recognition in the Wild

NeurIPS 2025poster

To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the fi…

Cited by 0SourceScholar
2025

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation

CVPR 2025poster

This paper presents Segment This Thing (STT), a new efficient image segmentation model designed to produce a single segment given a single point prompt. Instead of following prior work and increasing efficiency by decreasing model size, we gain efficiency by foveating input images. Given an image an…

2025

Sonata: Self-Supervised Learning of Reliable Point Representations

CVPR 2025highlight

In this paper, we question whether we have a reliable self-supervised point cloud model that can be used for diverse 3D tasks via simple linear probing, even with limited data and minimal computation. We find that existing 3D self-supervised learning approaches fall short when evaluated on represent…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

Nymeria: A Massive Collection of Egocentric Multi-modal Human Motion in the Wild

ECCV 2024poster

"We introduce - a large-scale, diverse, richly annotated human motion dataset collected in the wild with multiple multimodal egocentric devices. The dataset comes with a) full-body ground-truth motion; b) multiple multimodal egocentric data from Project Aria devices with videos, eye tracking, IMUs a…

2024

SceneScript: Reconstructing Scenes With An Autoregressive Structured Language Model

ECCV 2024poster

"We introduce , a method that directly produces full scene models as a sequence of structured language commands using an autoregressive, token-based approach. Our proposed scene representation is inspired by recent successes in transformers & LLMs, and departs from more traditional methods which com…

Cited by 25SourcePDFScholar
2023

Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception

ICCV 2023poster

We introduce the Aria Digital Twin (ADT) - an egocentric dataset captured using Aria glasses with extensive object, environment, and human level ground truth. This ADT release contains 200 sequences of real-world activities conducted by Aria wearers in two real indoor scenes with 398 object instance…

Cited by 52PDFcodeScholar
2023

Ego-Humans: An Ego-Centric 3D Multi-Human Benchmark

ICCV 2023oral

We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing egocentric benchmarks either capture single subject or indoor-only scenarios, which limit the generalization of computer vision algorithms…

Cited by 39PDFScholar
2023

OrienterNet: Visual Localization in 2D Public Maps With Neural Matching

CVPR 2023poster

Humans can orient themselves in their 3D environments using simple 2D maps. Differently, algorithms for visual localization mostly rely on complex 3D point clouds that are expensive to build, store, and maintain over time. We bridge this gap by introducing OrienterNet, the first deep neural network…

2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

LISA: Learning Implicit Shape and Appearance of Hands

CVPR 2022poster

This paper proposes a do-it-all neural model of human hands, named LISA. The model can capture accurate hand shape and appearance, generalize to arbitrary hand subjects, provide dense surface correspondences, be reconstructed from images in the wild and easily animated. We train LISA by minimizing t…

Cited by 80PDFScholar
2022

Neural 3D Video Synthesis From Multi-View Video

CVPR 2022oral

We propose a novel approach for 3D video synthesis that is able to represent multi-view video recordings of a dynamic real-world scene in a compact, yet expressive representation that enables high-quality view synthesis and motion interpolation. Our approach takes the high quality and compactness of…

Cited by 486PDFcodeScholar
2022

Self-Supervised Neural Articulated Shape and Appearance Models

CVPR 2022poster

Learning geometry, motion, and appearance priors of object classes is important for the solution of a large variety of computer vision problems. While the majority of approaches has focused on static objects, dynamic objects, especially with controllable articulation, are less explored. We propose a…

Cited by 39PDFcodeScholar
2021

ODAM: Object Detection, Association, and Mapping Using Posed RGB Video

ICCV 2021poster

Localizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system…

Cited by 34PDFcodeScholar
2020

Deep Local Shapes: Learning Local SDF Priors for Detailed 3D Reconstruction

ECCV 2020poster

Efficiently reconstructing complex and intricate surfaces at scale is a long-standing goal in machine perception. To address this problem we introduce Deep Local Shapes (DeepLS), a deep shape representation that enables high-quality 3D shape representation without prohibitive memory requirements. De…

Cited by 537SourcePDFScholar
2019

DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation

CVPR 2019oral

Computer graphics, 3D computer vision and robotics communities have produced multiple approaches to representing 3D geometry for rendering and reconstruction. These provide trade-offs across fidelity, efficiency and compression capabilities. In this work, we introduce DeepSDF, a learned continuous S…

Cited by 4353PDFScholar
2015

Depth-based tracking with physical constraints for robot manipulation

ICRA 2015poster

This work integrates visual and physical constraints to perform real-time depth-only tracking of articulated objects, with a focus on tracking a robot's manipulators and manipulation targets in realistic scenarios. As such, we extend DART, an existing visual articulated object tracker, to additional…

Cited by 83SourceScholar