← Search

Lourdes Agapito

27 accepted papers

2026

Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction

ICLR 2026poster

We address the 3D reconstruction of human faces from a single RGB image. To this end, we propose Pixel3DMM, a set of highly-generalized vision transformers which predict per-pixel geometric cues in order to constrain the optimization of a 3D morphable face model (3DMM). We exploit the latent feature…

Cited by 0SourcecodeScholar
2025

BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View Synthesis

ICCV 2025poster

We present billboard Splatting (BBSplat) - a novel approach for novel view synthesis based on textured geometric primitives. BBSplat represents the scene as a set of optimizable textured planar primitives with learnable RGB textures and alpha-maps to control their shape. BBSplat primitives can be us…

2025

Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects

CVPR 2025poster

Rapid advancements in text-to-3D generation require robust and scalable evaluation metrics that align closely with human judgment, a need unmet by current metrics such as PSNR and CLIP, which require ground-truth data or focus only on prompt fidelity. To address this, we introduce Gen3DEval, a novel…

2025

Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors

CVPR 2025poster

We present Pow3R, a novel large 3D vision regression model that is highly versatile in the input modalities it accepts. Unlike previous feed-forward models that lack any mechanism to exploit known camera or scene priors at test time, Pow3R incorporates any combination of auxiliary information such a…

Cited by 2SourcePDFScholar
2025

Semantic Cross-Pose Correspondence from a Single Example

ICRA 2025

This article focuses on predicting how an object can be transformed to a semantically meaningful pose relative to another object, given only one or few examples. Current pose correspondence methods rely on vast 3D object datasets and do not actively consider semantic information, which limits the ob

Cited by 1SourceScholar
2024

MonoNPHM: Dynamic Head Reconstruction from Monocular Videos

CVPR 2024highlight

We present Monocular Neural Parametric Head Models (MonoNPHM) for dynamic 3D head reconstructions from monocular RGB videos. To this end we propose a latent appearance space that parameterizes a texture field on top of a neural parametric model. We constrain predicted color values to be correlated w…

Cited by 19SourcePDFScholar
2024

MorpheuS: Neural Dynamic 360deg Surface Reconstruction from Monocular RGB-D Video

CVPR 2024poster

Neural rendering has demonstrated remarkable success in dynamic scene reconstruction. Thanks to the expressiveness of neural representations prior works can accurately capture the motion and achieve high-fidelity reconstruction of the target object. Despite this real-world video scenarios often feat…

Cited by 2SourcePDFScholar
2024

NViST: In the Wild New View Synthesis from a Single Image with Transformers

CVPR 2024poster

We propose NViST a transformer-based model for efficient and generalizable novel-view synthesis from a single image for real-world scenes. In contrast to many methods that are trained on synthetic data object-centred scenarios or in a category-specific manner NViST is trained on MVImgNet a large-sca…

2024

RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation

ICRA 2024poster

For robots to be useful outside labs and specialized factories we need a way to teach them new useful behaviors quickly. Current approaches lack either the generality to onboard new tasks without task-specific engineering, or else lack the data-efficiency to do so in an amount of time that enables p…

Cited by 45SourceScholar
2023

Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAM

CVPR 2023poster

We present Co-SLAM, a neural RGB-D SLAM system based on a hybrid representation, that performs robust camera tracking and high-fidelity surface reconstruction in real time. Co-SLAM represents the scene as a multi-resolution hash-grid to exploit its high convergence speed and ability to represent hig…

2023

Learning Neural Parametric Head Models

CVPR 2023poster

We propose a novel 3D morphable model for complete human heads based on hybrid neural fields. At the core of our model lies a neural parametric representation that disentangles identity and expressions in disjoint latent spaces. To this end, we capture a person's identity in a canonical space as a s…

Cited by 55SourcePDFScholar
2023

SeMLaPS: Real-Time Semantic Mapping With Latent Prior Networks and Quasi-Planar Segmentation

RA-L 2023

The availability of real-time semantics greatly improves the core geometric functionality of SLAM systems, enabling numerous robotic and AR/VR applications. We present a new methodology for real-time semantic mapping from RGB-D sequences that combines a 2D neural network and a 3D network based on a

Cited by 9SourcecodeScholar
2022

Few-Shot Keypoint Detection as Task Adaptation via Latent Embeddings

ICRA 2022poster

Dense object tracking, the ability to localize specific object points with pixel-level accuracy, is an important computer vision task with numerous downstream applications in robotics. Existing approaches either compute dense keypoint embeddings in a single forward pass, meaning the model is trained…

Cited by 3SourceScholar
2022

One-Shot Transfer of Affordance Regions? AffCorrs!

CoRL 2022poster

In this work, we tackle one-shot visual search of object parts. Given a single reference image of an object with annotated affordance regions, we segment semantically corresponding parts within a target scene. We propose AffCorrs, an unsupervised model that combines the properties of pre-trained D…

Cited by 44SourcecodeScholar
2020

S3K: Self-Supervised Semantic Keypoints for Robotic Manipulation via Multi-View Consistency

CoRL 2020

A robot’s ability to act is fundamentally constrained by what it can perceive. Many existing approaches to visual representation learning utilize general-purpose training criteria, e.g. image reconstruction, smoothness in latent space, or usefulness for control, or else make use of large datasets an

Cited by 0SourcePDFScholar
2018

Detect Globally, Label Locally: Learning Accurate 6-DOF Object Pose Estimation by Joint Segmentation and Coordinate Regression

RA-L 2018

Coordinate regression has established itself as one of the most successful current trends in model-based 6 degree of freedom (6-DOF) object pose estimation from a single image. The underlying idea is to train a system that can regress the three-dimensional coordinates of an object, given an input RG

Cited by 15SourceScholar
2018

DiverseNet: When One Right Answer Is Not Enough

CVPR 2018poster

Many structured prediction tasks in machine vision have a collection of acceptable answers, instead of one definitive ground truth answer. Segmentation of images, for example, is subject to human labeling bias. Similarly, there are multiple possible pixel values that could plausibly complete occlude…

Cited by 38SourcePDFScholar
2018

Structured Uncertainty Prediction Networks

CVPR 2018poster

This paper is the first work to propose a network to predict a structured uncertainty distribution for a synthesized image. Previous approaches have been mostly limited to predicting diagonal covariance matrices. Our novel model learns to predict a full Gaussian covariance matrix for each reconstruc…

Cited by 81SourcePDFScholar
2015

Direct, Dense, and Deformable: Template-Based Non-Rigid 3D Reconstruction From RGB Video

ICCV 2015poster

In this paper we tackle the problem of capturing the dense, detailed 3D geometry of generic, complex non-rigid meshes using a single RGB-only commodity video camera and a direct approach. While robust and even real-time solutions exist to this problem if the observed scene is static, for non-rigid…

Cited by 115PDFScholar
2015

Part-Based Modelling of Compound Scenes From Images

CVPR 2015poster

We propose a method to recover the structure of a compound scene from multiple silhouettes. Structure is expressed as a collection of 3D primitives chosen from a pre-defined library, each with an associated pose. This has several advantages over a volume or mesh representation both for estimation an…

Cited by 33SourcePDFScholar