← Search

Fernando De La Torre

39 accepted papers

2026

CGHair: Compact Gaussian Hair Reconstruction with Card Clustering

CVPR 2026

We present a compact pipeline for high-fidelity hair reconstruction from multi-view images. While recent 3D Gaussian Splatting (3DGS) methods achieve realistic results, they often require millions of primitives, leading to high storage and rendering costs. Observing that hair exhibits structural and

Cited by 0SourceScholar
2025

Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision

ICCV 2025poster

Detecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state-of-the-art (SOTA) results for this application. However, a significant challenge arises when models trained on data fro…

2025

GAS: Generative Avatar Synthesis from a Single Image

ICCV 2025poster

We present a unified and generalizable framework for synthesizing view-consistent and temporally coherent avatars from a single image, addressing the challenging task of single-image avatar generation. Existing diffusion-based methods often condition on sparse human templates (e.g., depth or normal…

2025

Improving Noise Efficiency in Privacy-preserving Dataset Distillation

ICCV 2025poster

Modern machine learning models heavily rely on large datasets that often include sensitive and private information, raising serious privacy concerns. Differentially private (DP) data generation offers a solution by creating synthetic datasets that limit the leakage of private information within a pr…

2025

LightSwitch: Multi-view Relighting with Material-guided Diffusion

ICCV 2025poster

Recent approaches for 3D relighting have shown promise in integrating 2D image relighting generative priors to alter the appearance of a 3D representation while preserving the underlying structure. Nevertheless, generative priors used for 2D relighting that directly relight from an input image do no…

Cited by 0SourcePDFScholar
2025

On the Fine-Grained Planning Abilities of VLM Web Agents

EMNLP 2025

Vision-Language Models (VLMs) have shown promise as web agents, yet their planning—the ability to devise strategies or action sequences to complete tasks—remains understudied. While prior works focus on VLM’s perception and overall success rates (i.e., goal completion), fine-grained investigation of

Cited by 0SourcePDFScholar
2024

Domain Gap Embeddings for Generative Dataset Augmentation

CVPR 2024poster

The performance of deep learning models is intrinsically tied to the quality volume and relevance of their training data. Gathering ample data for production scenarios often demands significant time and resources. Among various strategies data augmentation circumvents exhaustive data collection by g…

Cited by 7SourcePDFScholar
2024

Doubly Hierarchical Geometric Representations for Strand-based Human Hairstyle Generation

NeurIPS 2024poster

We introduce a doubly hierarchical generative representation for strand-based 3D hairstyle geometry that progresses from coarse, low-pass filtered guide hair to densely populated hair strands rich in high-frequency details. We employ the Discrete Cosine Transform (DCT) to separate low-frequency stru…

Cited by 0SourcePDFScholar
2024

FAMOUS: High-Fidelity Monocular 3D Human Digitization Using View Synthesis

ECCV 2024poster

"The advancement in deep implicit modeling and articulated models has significantly enhanced the process of digitizing human figures in 3D from just a single image. While state-of-the-art methods have greatly improved geometric precision, the challenge of accurately inferring texture remains, partic…

2024

Generalizable Human Gaussians for Sparse View Synthesis

ECCV 2024poster

"Recent progress in neural rendering has brought forth pioneering methods, such as NeRF and Gaussian Splatting, which revolutionize view rendering across various domains like AR/VR, gaming, and content creation. While these methods excel at interpolating within the training data, the challenge of ge…

2024

Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mamba

NeurIPS 2024poster

3D Hand reconstruction from a single RGB image is challenging due to the articulated motion, self-occlusion, and interaction with objects. Existing SOTA methods employ attention-based transformers to learn the 3D hand pose and shape, yet they do not fully achieve robust and accurate performance, pri…

2024

POET: Prompt Offset Tuning for Continual Human Action Adaptation

ECCV 2024oral

"As extended reality (XR) is redefining how users interact with computing devices, research in human action recognition is gaining prominence. Typically, models deployed on immersive computing devices are static and limited to their default set of classes. The goal of our research is to provide user…

2024

Visual Data Diagnosis and Debiasing with Concept Graphs

NeurIPS 2024poster

The widespread success of deep learning models today is owed to the curation of extensive datasets significant in size and complexity. However, such models frequently pick up inherent biases in the data during the training process, leading to unreliable predictions. Diagnosing and debiasing datasets…

2023

A Latent Space of Stochastic Diffusion Models for Zero-Shot Image Editing and Guidance

ICCV 2023poster

Diffusion models generate images by iterative denoising. Recent work has shown that by making the denoising process deterministic, one can encode real images into latent codes of the same size, which can be used for image editing. This paper explores the possibility of defining a latent space even w…

Cited by 98PDFcodeScholar
2023

Data-Free Class-Incremental Hand Gesture Recognition

ICCV 2023poster

This paper investigates data-free class-incremental learning (DFCIL) for hand gesture recognition from 3D skeleton sequences. In this class-incremental learning (CIL) setting, while incrementally registering the new classes, we do not have access to the training samples (i.e. data-free) of t…

Cited by 9PDFcodeScholar
2023

ITI-GEN: Inclusive Text-to-Image Generation

ICCV 2023oral

Text-to-image generative models often reflect the biases of the training data, leading to unequal representations of underrepresented groups. This study investigates inclusive text-to-image generative models that generate images based on human-written prompts and ensure the resulting images are unif…

Cited by 66PDFcodeScholar
2023

PATMAT: Person Aware Tuning of Mask-Aware Transformer for Face Inpainting

ICCV 2023poster

Generative models such as StyleGAN2 and Stable Diffusion have achieved state-of-the-art performance in computer vision tasks such as image synthesis, inpainting, and de-noising. However, current generative models for face inpainting often fail to preserve fine facial details and the identity of the…

Cited by 3PDFcodeScholar
2022

Generative Visual Prompt: Unifying Distributional Control of Pre-Trained Generative Models

NeurIPS 2022accept

Generative models (e.g., GANs, diffusion models) learn the underlying data distribution in an unsupervised manner. However, many applications of interest require sampling from a particular region of the output space or sampling evenly over a range of characteristics. For efficient sampling in these…

2022

Robust Egocentric Photo-Realistic Facial Expression Transfer for Virtual Reality

CVPR 2022poster

Social presence, the feeling of being there with a "real" person, will fuel the next generation of communication systems driven by digital humans in virtual reality (VR). The best 3D video-realistic VR avatars that minimize the uncanny effect rely on person-specific (PS) models. However, these PS mo…

Cited by 14PDFScholar
2021

High-Fidelity Face Tracking for AR/VR via Deep Lighting Adaptation

CVPR 2021poster

3D video avatars can empower virtual communications by providing compression, privacy, entertainment, and a sense of presence in AR/VR. Best 3D photo-realistic AR/VR avatars driven by video, that can minimize uncanny effects, rely on person-specific models. However, existing person-specific photo-re…

Cited by 29PDFScholar
2021

Implicit HRTF Modeling Using Temporal Convolutional Networks

ICASSP 2021accepted

Estimation of accurate head-related transfer functions (HRTFs) is crucial to achieve realistic binaural acoustic experiences. HRTFs depend on source/listener locations and are therefore expensive and cumbersome to measure; traditional approaches require listener-dependent measurements of HRTFs at th…

Cited by 0SourceScholar
2021

MeshTalk: 3D Face Animation From Speech Using Cross-Modality Disentanglement

ICCV 2021poster

This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible co-articulation or rely on person-specific models that limit their sca…

Cited by 241PDFcodeScholar
2020

3D Human Shape and Pose from a Single Low-Resolution Image with Self-Supervised Learning

ECCV 2020poster

3D human shape and pose estimation from monocular images has been an active area of research in computer vision, having a substantial impact on the development of new applications, from activity recognition to creating virtual avatars. Existing deep learning methods for 3D human shape and pose estim…

2020

Expressive Telepresence via Modular Codec Avatars

ECCV 2020poster

VR telepresence consists of interacting with another human in a virtual space represented by an avatar. Today most avatars are cartoon-like, but soon the technology will allow video-realistic ones. This paper aims in this direction and presents Modular Codec Avatars (MCA), a method to generate hyper…

Cited by 40SourcePDFScholar
2018

Inverse Composition Discriminative Optimization for Point Cloud Registration

CVPR 2018poster

Rigid Point Cloud Registration (PCReg) refers to the problem of finding the rigid transformation between two sets of point clouds. This problem is particularly important due to the advances in new 3D sensing hardware, and it is challenging because neither the correspondence nor the transformation pa…

Cited by 24SourcePDFScholar
2017

Discriminative Optimization: Theory and Applications to Point Cloud Registration

CVPR 2017poster

Many computer vision problems are formulated as the optimization of a cost function. This approach faces two main challenges: (1) designing a cost function with a local optimum at an acceptable solution, and (2) developing an efficient numerical method to search for one (or multiple) of these local…

Cited by 42PDFScholar
2016

Motion From Structure (MfS): Searching for 3D Objects in Cluttered Point Trajectories

CVPR 2016spotlight

Object detection has been a long standing problem in computer vision, and state-of-the-art approaches rely on the use of sophisticated features and/or classifiers. However, these learning-based approaches heavily depend on the quality and quantity of labeled data, and do not generalize well to extre…

Cited by 4PDFScholar
2015

Confidence Preserving Machine for Facial Action Unit Detection

ICCV 2015poster

Varied sources of error contribute to the challenge of facial action unit detection. Previous approaches address specific and known sources. However, many sources are unknown. To address the ubiquity of error, we propose a Confident Preserving Machine (CPM) that follows an easy-to-hard classificatio…

Cited by 82PDFScholar
2015

Joint Patch and Multi-Label Learning for Facial Action Unit Detection

CVPR 2015poster

The face is one of the most powerful channel of non-verbal communication. The most commonly used taxonomy to describe facial behaviour is the Facial Action Coding System (FACS). FACS segments the visible effects of facial muscle activation into 30+ action units (AUs). AUs, which may occur alone an…

Cited by 246SourcePDFScholar
2015

Unsupervised Synchrony Discovery in Human Interaction

ICCV 2015poster

People are inherently social. Social interaction plays an important and natural role in human behavior. Most computational methods focus on individuals alone rather than in social context. They also require labelled training data. We present an unsupervised approach to discover interpersonal synchro…

Cited by 19PDFScholar