← Search

Pablo Garrido

10 accepted papers

2025

KinMo: Kinematic-aware Human Motion Understanding and Generation

ICCV 2025poster

Current human motion synthesis frameworks rely on global action descriptions, creating a modality gap that limits both motion understanding and generation capabilities. A single coarse description, such as "run", fails to capture essential details like variations in speed, limb positioning, and kine…

2025

VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing

CVPR 2025poster

We present an implicit video representation for occlusions, appearance, and motion disentanglement from monocular videos, which we refer to as Video Spatiotemporal Splines (VideoSPatS).Unlike previous methods that map time and coordinates to deformation and canonical colors, our VideoSPatS maps inpu…

2023

Few-Shot Geometry-Aware Keypoint Localization

CVPR 2023poster

Supervised keypoint localization methods rely on large manually labeled image datasets, where objects can deform, articulate, or occlude. However, creating such large keypoint labels is time-consuming and costly, and is often error-prone due to inconsistent labeling. Thus, we desire an approach that…

2023

Implicit Neural Head Synthesis via Controllable Local Deformation Fields

CVPR 2023poster

High-quality reconstruction of controllable 3D head avatars from 2D videos is highly desirable for virtual human applications in movies, games, and telepresence. Neural implicit fields provide a powerful representation to model 3D head avatars with personalized shape, expressions, and facial parts,…

Cited by 12SourcePDFScholar
2023

Unsupervised Facial Performance Editing via Vector-Quantized StyleGAN Representations

ICCV 2023poster

High-fidelity virtual human avatar applications create a need for photorealistic video face synthesis with controllable semantic editing over facial features. While recent generative neural methods have shown significant progress in portrait video synthesis, intuitive facial control, e.g., of mouth…

Cited by 1PDFcodeScholar
2019

FML: Face Model Learning From Videos

CVPR 2019oral

Monocular image-based 3D reconstruction of faces is a long-standing problem in computer vision. Since image data is a 2D projection of a 3D face, the resulting depth ambiguity makes the problem ill-posed. Most existing methods rely on data-driven priors that are built from limited 3D face scans. In…

Cited by 179PDFScholar
2018

Self-Supervised Multi-Level Face Model Learning for Monocular Reconstruction at Over 250 Hz

CVPR 2018poster

The reconstruction of dense 3D models of face geometry and appearance from a single image is highly challenging and ill-posed. To constrain the problem, many approaches rely on strong priors, such as parametric face models learned from limited 3D scan data. However, prior models restrict generalizat…

Cited by 308SourcePDFScholar
2017

MoFA: Model-Based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction

ICCV 2017oral

In this work we propose a novel model-based deep convolutional autoencoder that addresses the highly challenging problem of reconstructing a 3D human face from a single in-the-wild color image. To this end, we combine a convolutional encoder network with an expert-designed generative model that serv…

Cited by 688PDFScholar