← Search

David Picard

12 accepted papers

2026

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

ICML 2026poster

The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward,…

Cited by 0SourceScholar
2026

POLYNOMIAL MIXING FOR EFFICIENT SELF-SUPERVISED SPEECH ENCODERS

ICASSP 2026oral

State-of-the-art speech-to-text models typically employ Transformer-based encoders that model token dependencies via self-attention mechanisms. However, the quadratic complexity of self-attention in both memory and computation imposes significant constraints on scalability. In this work, we propose…

Cited by 0SourcePDFScholar
2026

Random Process Flow Matching: Generative Implicit Representations of Multivariate Random Fields

ICML 2026poster

Generative modeling provides a powerful framework for learning data distributions. These models initially relied on probabilistic methods such as Gaussian Processes (GP) for uncertainty-aware predictions and shifted towards larger trainable models to learn more complex distributions. In this work, w…

Cited by 0SourceScholar
2025

Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation

CVPR 2025poster

Global visual geolocation predicts where an image was captured on Earth. Since images vary in how precisely they can be localized, this task inherently involves a significant degree of ambiguity. However, existing approaches are deterministic and overlook this aspect. In this paper, we aim to close…

2024

Don't Drop Your Samples! Coherence-Aware Training Benefits Conditional Diffusion

CVPR 2024highlight

Conditional diffusion models are powerful generative models that can leverage various types of conditional information such as class labels segmentation masks or text captions. However in many real-world scenarios conditional information may be noisy or unreliable due to human annotation errors or w…

Cited by 3SourcePDFScholar
2023

Unveiling the Latent Space Geometry of Push-Forward Generative Models

ICML 2023poster

Many deep generative models are defined as a push-forward of a Gaussian measure by a continuous generator, such as Generative Adversarial Networks (GANs) or Variational Auto-Encoders (VAEs). This work explores the latent space of such deep generative models. A key issue with these models is their te…

Cited by 4SourcePDFScholar
2022

SCAM! Transferring Humans between Images with Semantic Cross Attention Modulation

ECCV 2022poster

"A large body of recent work targets semantically conditioned image generation. Most such methods focus on the narrower task of pose transfer and ignore the more challenging task of subject transfer that consists in not only transferring the pose but also the appearance and background. In this work,…

2021

Triggering Failures: Out-of-Distribution Detection by Learning From Local Adversarial Attacks in Semantic Segmentation

ICCV 2021poster

In this paper, we tackle the detection of out-of-distribution (OOD) objects in semantic segmentation. By analyzing the literature, we found that current methods are either accurate or fast but not both which limits their usability in real world applications. To get the best of both aspects, we propo…

Cited by 58PDFcodeScholar
2019

Metric Learning With HORDE: High-Order Regularizer for Deep Embeddings

ICCV 2019poster

Learning an effective similarity measure between image representations is key to the success of recent advances in visual search tasks (e.g. verification or zero-shot learning). Although the metric learning part is well addressed, this metric is usually computed over the average of the extracted dee…

Cited by 84PDFcodeScholar
2018

2D/3D Pose Estimation and Action Recognition Using Multitask Deep Learning

CVPR 2018poster

Action recognition and human pose estimation are closely related but both problems are generally handled as distinct tasks in the literature. In this work, we propose a multitask framework for jointly 2D and 3D pose estimation from still images and human action recognition from video sequences. We s…

Cited by 678SourcePDFScholar
2018

Image Reassembly Combining Deep Learning and Shortest Path Problem

ECCV 2018poster

This paper addresses the problem of reassembling images from disjointed fragments. More specifically, given an unordered set of fragments, we aim at reassembling one or several possibly incomplete images. The main contributions of this work are: 1) several deep neural architectures to predict the re…

Cited by 42SourcePDFScholar