← Search

Patrick Perez

29 accepted papers

2026

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

ICML 2026poster

We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference time. Traditional language modeling is an intrinsically ill-posed problem: given a context, multiple ``futures'' may be…

Cited by 0SourceScholar
2026

Understanding Data Temporality Impact on Large Language Models Pre-training

ICML 2026poster

Large language models (LLMs) are typically trained on shuffled corpora, yielding models whose knowledge is frozen at training time and whose temporal grounding remains poorly understood. In this work, we study the impact of pretraining dynamics on the acquisition of time-sensitive factual knowledge,…

Cited by 0SourceScholar
2025

ToddlerDiffusion: Interactive Structured Image Generation with Cascaded Schrödinger Bridge

ICLR 2025poster

Diffusion models break down the challenging task of generating data from high-dimensional distributions into a series of easier denoising steps. Inspired by this paradigm, we propose a novel approach that extends the diffusion framework into modality space, decomposing the complex task of RGB image…

2024

ManiPose: Manifold-Constrained Multi-Hypothesis 3D Human Pose Estimation

NeurIPS 2024poster

We propose ManiPose, a manifold-constrained multi-hypothesis model for human-pose 2D-to-3D lifting. We provide theoretical and empirical evidence that, due to the depth ambiguity inherent to monocular 3D human pose estimation, traditional regression models suffer from pose-topology consistency issue…

2024

Winner-takes-all learners are geometry-aware conditional density estimators

ICML 2024poster

Winner-takes-all training is a simple learning paradigm, which handles ambiguous tasks by predicting a set of plausible hypotheses. Recently, a connection was established between Winner-takes-all training and centroidal Voronoi tessellations, showing that, once trained, hypotheses should quantize op…

2023

POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images

NeurIPS 2023poster

We describe an approach to predict open-vocabulary 3D semantic voxel occupancy map from input 2D images with the objective of enabling 3D grounding, segmentation and retrieval of free-form language queries. This is a challenging problem because of the 2D-3D ambiguity and the open-vocabulary nature o…

2023

Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis

NeurIPS 2023poster

We introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may be sampled for each training input. Multiple Choice Learning is a simple framework to tackle multimodal density estimatio…

2023

Self-supervised learning with rotation-invariant kernels

ICLR 2023top-25%

We introduce a regularization loss based on kernel mean embeddings with rotation-invariant kernels on the hypersphere (also known as dot-product kernels) for self-supervised learning of image representations. Besides being fully competitive with the state of the art, our method significantly reduces…

2022

LaRa: Latents and Rays for Multi-Camera Bird’s-Eye-View Semantic Segmentation

CoRL 2022poster

Recent works in autonomous driving have widely adopted the bird’seye-view (BEV) semantic map as an intermediate representation of the world. Online prediction of these BEV maps involves non-trivial operations such as multi-camera data extraction as well as fusion and projection into a common topview…

Cited by 40SourcecodeScholar
2021

Large-Scale Unsupervised Object Discovery

NeurIPS 2021poster

Existing approaches to unsupervised object discovery (UOD) do not scale up to large datasets without approximations that compromise their performance. We propose a novel formulation of UOD as a ranking problem, amenable to the arsenal of distributed methods available for eigenvalue problems and link…

2021

OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning

CVPR 2021poster

Learning image representations without human supervision is an important and active research field. Several recent approaches have successfully leveraged the idea of making such a representation invariant under different types of perturbations, especially via contrastive-based instance discriminatio…

Cited by 125PDFcodeScholar
2021

Semantic Palette: Guiding Scene Generation With Class Proportions

CVPR 2021poster

Despite the recent progress of generative adversarial networks (GANs) at synthesizing photo-realistic images, producing complex urban scenes remains a challenging problem. Previous works break down scene generation into two consecutive phases: unconditional semantic layout synthesis and image synthe…

Cited by 18PDFcodeScholar
2020

Learning Representations by Predicting Bags of Visual Words

CVPR 2020poster

Self-supervised representation learning targets to learn convnet-based image representations from unlabeled data. Inspired by the success of NLP methods in this area, in this work we propose a self-supervised approach based on spatially dense image descriptions that encode discrete visual concepts,…

Cited by 132PDFcodeScholar
2020

StyleRig: Rigging StyleGAN for 3D Control Over Portrait Images

CVPR 2020oral

StyleGAN generates photorealistic portrait images of faces with eyes, teeth, hair and context (neck, shoulders, background), but lacks a rig-like control over semantic face parameters that are interpretable in 3D, such as face pose, expressions, and scene illumination. Three-dimensional morphable fa…

Cited by 473PDFScholar
2020

xMUDA: Cross-Modal Unsupervised Domain Adaptation for 3D Semantic Segmentation

CVPR 2020poster

Unsupervised Domain Adaptation (UDA) is crucial to tackle the lack of annotations in a new domain. There are many multi-modal datasets, but most UDA approaches are uni-modal. In this work, we explore how to learn from multi-modality and propose cross-modal UDA (xMUDA) where we assume the presence of…

Cited by 224PDFcodeScholar
2019

ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation

CVPR 2019oral

Semantic segmentation is a key problem for many computer vision tasks. While approaches based on convolutional neural networks constantly break new records on different benchmarks, generalizing well to diverse testing environments remains a major challenge. In numerous real-world applications, there…

Cited by 1726PDFcodeScholar
2019

Boosting Few-Shot Visual Learning With Self-Supervision

ICCV 2019poster

Few-shot learning and self-supervised learning address different facets of the same problem: how to train a model with little or no labeled data. Few-shot learning aims for optimization methods and models that can learn efficiently to recognize patterns in the low data regime. Self-supervised learni…

Cited by 511PDFcodeScholar
2019

DADA: Depth-Aware Domain Adaptation in Semantic Segmentation

ICCV 2019poster

Unsupervised domain adaptation (UDA) is important for applications where large scale annotation of representative data is challenging. For semantic segmentation in particular, it helps deploy on real "target domain" data models that are trained on annotated images from a different "source domain", n…

Cited by 263PDFcodeScholar
2019

FML: Face Model Learning From Videos

CVPR 2019oral

Monocular image-based 3D reconstruction of faces is a long-standing problem in computer vision. Since image data is a 2D projection of a 3D face, the resulting depth ambiguity makes the problem ill-posed. Most existing methods rely on data-driven priors that are built from limited 3D face scans. In…

Cited by 179PDFScholar
2019

SoDeep: A Sorting Deep Net to Learn Ranking Loss Surrogates

CVPR 2019oral

Several tasks in machine learning are evaluated using non-differentiable metrics such as mean average precision or Spearman correlation. However, their non-differentiability prevents from using them as objective functions in a learning framework. Surrogate and relaxation methods exist but tend to be…

Cited by 90PDFcodeScholar
2019

Unsupervised Image Matching and Object Discovery as Optimization

CVPR 2019poster

Learning with complete or partial supervision is power- ful but relies on ever-growing human annotation efforts. As a way to mitigate this serious problem, as well as to serve specific applications, unsupervised learning has emerged as an important field of research. In computer vision, unsu- pervis…

Cited by 77PDFcodeScholar
2019

WoodScape: A Multi-Task, Multi-Camera Fisheye Dataset for Autonomous Driving

ICCV 2019oral

Fisheye cameras are commonly employed for obtaining a large field of view in surveillance, augmented reality and in particular automotive applications. In spite of their prevalence, there are few public datasets for detailed evaluation of computer vision algorithms on fisheye images. We release the…

Cited by 350PDFcodeScholar
2017

Kernel Square-Loss Exemplar Machines for Image Retrieval

CVPR 2017poster

Zepeda and Perez have recently demonstrated the promise of the exemplar SVM (ESVM) as a feature encoder for image retrieval. This paper extends this approach in several directions: We first show that replacing the hinge loss by the square loss in the ESVM cost function significantly reduces encoding…

Cited by 13PDFScholar
2017

MoFA: Model-Based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction

ICCV 2017oral

In this work we propose a novel model-based deep convolutional autoencoder that addresses the highly challenging problem of reconstructing a 3D human face from a single in-the-wild color image. To this end, we combine a convolutional encoder network with an expert-designed generative model that serv…

Cited by 688PDFScholar
2017

ROAM: A Rich Object Appearance Model With Application to Rotoscoping

CVPR 2017poster

Rotoscoping, the detailed delineation of scene elements through a video shot, is a painstaking task of tremendous importance in professional post-production pipelines. While pixel-wise segmentation techniques can help for this task, professional rotoscoping tools rely on parametric curves that offer…

Cited by 6PDFScholar
2017

SUBIC: A Supervised, Structured Binary Code for Image Search

ICCV 2017spotlight

For large-scale visual search, highly compressed yet meaningful representations of images are essential. Structured vector quantizers based on product quantization and its variants are usually employed to achieve such compression while minimizing the loss of accuracy. Yet, unlike binary hashing sche…

Cited by 102PDFScholar
2016

Determining Occlusions From Space and Time Image Reconstructions

CVPR 2016poster

The problem of localizing occlusions between consecutive frames of a video is important but rarely tackled on its own. In most works, it is tightly interleaved with the computation of accurate optical flows, which leads to a delicate chicken-and-egg problem. With this in mind, we propose a novel app…

Cited by 13PDFScholar