← Search

Albert Pumarola

20 accepted papers

2025

Adaptive Guidance: Training-free Acceleration of Conditional Diffusion Models

AAAI 2025technical

This paper presents a comprehensive study on the role of Classifier-Free Guidance (CFG) in text-conditioned diffusion models from the perspective of inference efficiency. In particular, we relax the default choice of applying CFG in all diffusion steps and instead propose to search for more efficien…

Cited by 8SourcePDFScholar
2025

Autoregressive Distillation of Diffusion Transformers

CVPR 2025poster

Diffusion models with transformer architectures have demonstrated promising capabilities in generating high-fidelity images and scalability for high resolution. However, iterative sampling process required for synthesis is very resource-intensive. A line of work has focused on distilling solutions…

2025

FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute

CVPR 2025highlight

Despite their remarkable performance, modern Diffusion Transformers (DiTs) are hindered by substantial resource requirements during inference, stemming from the fixed and large amount of compute needed for each denoising step. In this work, we revisit the conventional static paradigm that allocates…

Cited by 1SourcePDFScholar
2025

Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment

ICLR 2025oral

The performance of large language models (LLMs) is closely linked to their underlying size, leading to ever-growing networks and hence slower inference. Speculative decoding has been proposed as a technique to accelerate autoregressive generation, leveraging a fast draft model to propose candidate t…

Cited by 1SourcePDFScholar
2024

Bespoke Non-Stationary Solvers for Fast Sampling of Diffusion and Flow Models

ICML 2024poster

This paper introduces Bespoke Non-Stationary (BNS) Solvers, a solver distillation approach to improve sample efficiency of Diffusion and Flow models. BNS solvers are based on a family of non-stationary solvers that provably subsumes existing numerical ODE solvers and consequently demonstrate conside…

Cited by 3SourcePDFScholar
2024

Bespoke Solvers for Generative Flow Models

ICLR 2024spotlight

Diffusion or flow-based models are powerful generative paradigms that are notoriously hard to sample as samples are defined as solutions to high-dimensional Ordinary or Stochastic Differential Equations (ODEs/SDEs) which require a large Number of Function Evaluations (NFE) to approximate well. Exist…

Cited by 20SourcePDFScholar
2023

Avatars Grow Legs: Generating Smooth Human Motion From Sparse Tracking Inputs With Diffusion Model

CVPR 2023poster

With the recent surge in popularity of AR/VR applications, realistic and accurate control of 3D full-body avatars has become a highly demanded feature. A particular challenge is that only a sparse tracking signal is available from standalone HMDs (Head Mounted Devices), often limited to tracking the…

2023

Re-ReND: Real-Time Rendering of NeRFs across Devices

ICCV 2023poster

This paper proposes a novel approach for rendering a pre-trained Neural Radiance Field (NeRF) in real-time on resource-constrained devices. We introduce Re-ReND, a method enabling Real-time Rendering of NeRFs across Devices. Re-ReND is designed to achieve real-time performance by converting the NeRF…

Cited by 21PDFcodeScholar
2022

VisCo Grids: Surface Reconstruction with Viscosity and Coarea Grids

NeurIPS 2022accept

Surface reconstruction has been seeing a lot of progress lately by utilizing Implicit Neural Representations (INRs). Despite their success, INRs often introduce hard to control inductive bias (i.e., the solution surface can exhibit unexplainable behaviours), have costly inference, and are slow to tr…

Cited by 19SourcePDFScholar
2021

D-NeRF: Neural Radiance Fields for Dynamic Scenes

CVPR 2021poster

Neural rendering techniques combining machine learning with geometric reasoning have arisen as one of the most promising approaches for synthesizing novel views of a scene from a sparse set of images. Among these, stands out the Neural radiance fields (NeRF), which trains a deep network to map 5D in…

Cited by 1585PDFScholar
2021

H3D-Net: Few-Shot High-Fidelity 3D Head Reconstruction

ICCV 2021poster

Recent learning approaches that implicitly represent surface geometry using coordinate-based neural representations have shown impressive results in the problem of multi-view 3D reconstruction. The effectiveness of these techniques is, however, subject to the availability of a large number (several…

Cited by 104PDFScholar
2021

SMPLicit: Topology-Aware Generative Model for Clothed People

CVPR 2021poster

In this paper we introduce SMPLicit, a novel generative model to jointly represent body pose, shape and clothing geometry. In contrast to existing learning-based approaches that require training specific models for each type of garment, SMPLicit can represent in a unified manner different garment to…

Cited by 219PDFcodeScholar
2020

C-Flow: Conditional Generative Flow Models for Images and 3D Point Clouds

CVPR 2020poster

Flow-based generative models have highly desirable properties like exact log-likelihood evaluation and exact latent-variable inference, however they are still in their infancy and have not received as much attention as alternative generative models. In this paper, we introduce C-Flow, a novel condit…

Cited by 113PDFScholar
2020

GanHand: Predicting Human Grasp Affordances in Multi-Object Scenes

CVPR 2020oral

The rise of deep learning has brought remarkable progress in estimating hand geometry from images where the hands are part of the scene. This paper focuses on a new problem not explored so far, consisting in predicting how a human would grasp one or several objects, given a single RGB image of these…

Cited by 201PDFScholar
2019

3DPeople: Modeling the Geometry of Dressed Humans

ICCV 2019poster

Recent advances in 3D human shape estimation build upon parametric representations that model very well the shape of the naked body, but are not appropriate to represent the clothing geometry. In this paper, we present an approach to model dressed humans and predict their geometry from single images…

Cited by 152PDFScholar
2018

GANimation: Anatomically-aware Facial Animation from a Single Image

ECCV 2018poster

Recent advances in Generative Adversarial Networks (GANs) have shown impressive results for task of facial expression synthesis. The most successful architecture is StarGAN, that conditions GANs' generation process with images of a specific domain, namely a set of images of persons sharing the same…

2018

Geometry-Aware Network for Non-Rigid Shape Prediction From a Single View

CVPR 2018poster

We propose a method for predicting the 3D shape of a deformable surface from a single view. By contrast with previous approaches, we do not need a pre-registered template of the surface, and our method is robust to the lack of texture and partial occlusions. At the core of our approach is a geometry…

Cited by 66SourcePDFScholar
2018

Unsupervised Person Image Synthesis in Arbitrary Poses

CVPR 2018poster

We present a novel approach for synthesizing photo-realistic images of people in arbitrary poses using generative adversarial learning. Given an input image of a person and a desired pose represented by a 2D skeleton, our model renders the image of the same person under the new pose, synthesizing no…

Cited by 211SourcePDFScholar
2017

PL-SLAM: Real-time monocular visual SLAM with points and lines

ICRA 2017poster

Low textured scenes are well known to be one of the main Achilles heels of geometric computer vision algorithms relying on point correspondences, and in particular for visual SLAM. Yet, there are many environments in which, despite being low textured, one can still reliably estimate line-based geome…

Cited by 551SourceScholar