← Search

Enric Corona

13 accepted papers

2025

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

CVPR 2025poster

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-to-3d-motion diffusion model, and 2) a novel diffusion-based architecture that a…

Cited by 27SourcePDFScholar
2024

DiffHuman: Probabilistic Photorealistic 3D Reconstruction of Humans

CVPR 2024poster

We present DiffHuman a probabilistic method for photorealistic 3D human reconstruction from a single RGB image. Despite the ill-posed nature of this problem most methods are deterministic and output a single solution often resulting in a lack of geometric detail and blurriness in unseen or uncertain…

Cited by 5SourcePDFScholar
2024

Instant 3D Human Avatar Generation using Image Diffusion Models

ECCV 2024poster

"We present , a method for fast, high quality 3D human avatar generation from different input modalities, such as images and text prompts and with control over the generated pose and shape. The common theme is the use of diffusion-based image generation networks that are specialized for each particu…

Cited by 5SourcePDFScholar
2023

Structured 3D Features for Reconstructing Controllable Avatars

CVPR 2023poster

We introduce Structured 3D Features, a model based on a novel implicit 3D representation that pools pixel-aligned image features onto dense 3D points sampled from a parametric, statistical human mesh surface. The 3D points have associated semantics and can move freely in 3D space. This allows for op…

Cited by 39SourcePDFScholar
2022

LISA: Learning Implicit Shape and Appearance of Hands

CVPR 2022poster

This paper proposes a do-it-all neural model of human hands, named LISA. The model can capture accurate hand shape and appearance, generalize to arbitrary hand subjects, provide dense surface correspondences, be reconstructed from images in the wild and easily animated. We train LISA by minimizing t…

Cited by 80PDFScholar
2022

Learned Vertex Descent: A New Direction for 3D Human Model Fitting

ECCV 2022poster

"We propose a novel optimization-based paradigm for 3D human shape fitting on images. In contrast to existing approaches that directly regress the parameters of a low-dimensional statistical body model (e.g. SMPL) from input images, we propose training a deep network that, given solely image feature…

Cited by 40SourcePDFScholar
2021

D-NeRF: Neural Radiance Fields for Dynamic Scenes

CVPR 2021poster

Neural rendering techniques combining machine learning with geometric reasoning have arisen as one of the most promising approaches for synthesizing novel views of a scene from a sparse set of images. Among these, stands out the Neural radiance fields (NeRF), which trains a deep network to map 5D in…

Cited by 1585PDFScholar
2021

Multi-FinGAN: Generative Coarse-To-Fine Sampling of Multi-Finger Grasps

ICRA 2021poster

While there exists many methods for manipulating rigid objects with parallel-jaw grippers, grasping with multi-finger robotic hands remains a quite unexplored research topic. Reasoning and planning collision-free trajectories on the additional degrees of freedom of several fingers represents an impo…

Cited by 63SourcecodeScholar
2021

SMPLicit: Topology-Aware Generative Model for Clothed People

CVPR 2021poster

In this paper we introduce SMPLicit, a novel generative model to jointly represent body pose, shape and clothing geometry. In contrast to existing learning-based approaches that require training specific models for each type of garment, SMPLicit can represent in a unified manner different garment to…

Cited by 219PDFcodeScholar
2020

GanHand: Predicting Human Grasp Affordances in Multi-Object Scenes

CVPR 2020oral

The rise of deep learning has brought remarkable progress in estimating hand geometry from images where the hands are part of the scene. This paper focuses on a new problem not explored so far, consisting in predicting how a human would grasp one or several objects, given a single RGB image of these…

Cited by 201PDFScholar