← Search

François Fleuret

30 accepted papers

2026

MODUS: Decoder-only Any-to-Any Modeling of Diverse Modalities

ICML 2026poster

Any-to-any modeling aims to flexibly relate arbitrary modalities within a single system, a requirement that arises across multimodal learning and scientific domains such as ecology and astronomy. However, existing any-to-any approaches are typically trained from scratch using encoder–decoder or diff…

Cited by 0SourceScholar
2026

Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising

ICML 2026poster

Normalization Equivariance (NE), equivariance to global contrast and brightness transforms, improves robustness to distribution shift in image-to-image prediction. Existing methods enforce this prior by constraining internal layers to NE-compatible families, limiting compatibility with standard comp…

Cited by 0SourceScholar
2025

LiNeS: Post-training Layer Scaling Prevents Forgetting and Enhances Model Merging

ICLR 2025poster

Fine-tuning pre-trained models has become the standard approach to endow them with specialized knowledge, but it poses fundamental challenges. In particular, (i) fine-tuning often leads to catastrophic forgetting, where improvements on a target domain degrade generalization on other tasks, and (ii)…

2025

Pareto Low-Rank Adapters: Efficient Multi-Task Learning with Preferences

ICLR 2025poster

Multi-task trade-offs in machine learning can be addressed via Pareto Front Learning (PFL) methods that parameterize the Pareto Front (PF) with a single model. PFL permits to select the desired operational point during inference, contrary to traditional Multi-Task Learning (MTL) that optimizes for a…

Cited by 6SourcePDFScholar
2024

DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging

NeurIPS 2024poster

The transformer architecture by Vaswani et al. (2017) is now ubiquitous across application domains, from natural language processing to speech processing and image understanding. We propose DenseFormer, a simple modification to the standard architecture that improves the perplexity of the model with…

Cited by 6SourcePDFScholar
2024

Diffusion for World Modeling: Visual Details Matter in Atari

NeurIPS 2024spotlight

World models constitute a promising approach for training reinforcement learning agents in a safe and sample-efficient manner. Recent world models predominantly operate on sequences of discrete latent variables to model environment dynamics. However, this compression into a compact discrete represen…

2024

Localizing Task Information for Improved Model Merging and Compression

ICML 2024poster

Model merging and task arithmetic have emerged as promising scalable approaches to merge multiple single-task checkpoints to one multi-task model, but their applicability is reduced by significant performance loss. Previous works have linked these drops to interference in the weight space and erasur…

2023

Agree to Disagree: Diversity through Disagreement for Better Transferability

ICLR 2023top-5%

Gradient-based learning algorithms have an implicit \emph{simplicity bias} which in effect can limit the diversity of predictors being sampled by the learning procedure. This behavior can hinder the transferability of trained models by (i) favoring the learning of simpler but spurious features --- p…

2023

ESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance Fields

CVPR 2023highlight

We present ESLAM, an efficient implicit neural representation method for Simultaneous Localization and Mapping (SLAM). ESLAM reads RGB-D frames with unknown camera poses in a sequential manner and incrementally reconstructs the scene representation while estimating the current camera position in the…

Cited by 177SourcePDFScholar
2023

Fast Attention Over Long Sequences With Dynamic Sparse Flash Attention

NeurIPS 2023poster

Transformer-based language models have found many diverse applications requiring them to process sequences of increasing length. For these applications, the causal self-attention---which is the only component scaling quadratically w.r.t. the sequence length---becomes a central concern. While many wo…

Cited by 12SourcePDFScholar
2023

Pareto Manifold Learning: Tackling multiple tasks via ensembles of single-task models

ICML 2023poster

In Multi-Task Learning (MTL), tasks may compete and limit the performance achieved on each other, rather than guiding the optimization to a solution, superior to all its single-task trained counterparts. Since there is often not a unique solution optimal for all tasks, practitioners have to balance…

2023

SUPA: A Lightweight Diagnostic Simulator for Machine Learning in Particle Physics

NeurIPS 2023poster

Deep learning methods have gained popularity in high energy physics for fast modeling of particle showers in detectors. Detailed simulation frameworks such as the gold standard \textsc{Geant4} are computationally intensive, and current deep generative architectures work on discretized, lower resolut…

Cited by 3SourcePDFScholar
2022

Efficient Training of Low-Curvature Neural Networks

NeurIPS 2022accept

Standard deep neural networks often have excess non-linearity, making them susceptible to issues such as low adversarial robustness and gradient instability. Common methods to address these downstream issues, such as adversarial training, are expensive and often sacrifice predictive accuracy. In…

Cited by 25SourcePDFScholar
2022

Flowification: Everything is a normalizing flow

NeurIPS 2022accept

The two key characteristics of a normalizing flow is that it is invertible (in particular, dimension preserving) and that it monitors the amount by which it changes the likelihood of data points as samples are propagated along the network. Recently, multiple generalizations of normalizing flows have…

2021

DepthInSpace: Exploitation and Fusion of Multiple Video Frames for Structured-Light Depth Estimation

ICCV 2021poster

We present DepthInSpace, a self-supervised deep-learning method for depth estimation using a structured-light camera. The design of this method is motivated by the commercial use case of embedded depth sensors in nowadays smartphones. We first propose to use estimated optical flow from ambient infor…

Cited by 14PDFScholar
2021

Taming GANs with Lookahead-Minmax

ICLR 2021poster

Generative Adversarial Networks are notoriously challenging to train. The underlying minmax optimization is highly susceptible to the variance of the stochastic gradient and the rotational component of the associated game vector field. To tackle these challenges, we propose the Lookahead algorithm f…

2020

Optimizer Benchmarking Needs to Account for Hyperparameter Tuning

ICML 2020poster

The performance of optimizers, particularly in deep learning, depends considerably on their chosen hyperparameter configuration. The efficacy of optimizers is often studied under near-optimal problem-specific hyperparameters, and finding these settings may be prohibitively costly for practitioners.…

Cited by 62SourcePDFScholar
2020

Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

ICML 2020poster

Transformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input’s length, they are prohibitively slow for very long sequences. To address this limitation, we express the self-attention as a linear dot-product of kernel feature maps and ma…

2019

Reducing Noise in GAN Training with Variance Reduced Extragradient

NeurIPS 2019poster

We study the effect of the stochastic gradient noise on the training of generative adversarial networks (GANs) and show that it can prevent the convergence of standard game optimization methods, while the batch version converges. We address this issue with a novel stochastic variance-reduced extragr…

Cited by 178SourcePDFScholar
2018

Practical Deep Stereo (PDS): Toward applications-friendly deep stereo matching

NeurIPS 2018poster

End-to-end deep-learning networks recently demonstrated extremely good performance for stereo matching. However, existing networks are difficult to use for practical applications since (1) they are memory-hungry and unable to process even modest-size images, (2) they have to be fully re-trained to h…

Cited by 149SourcePDFScholar
2018

WILDTRACK: A Multi-Camera HD Dataset for Dense Unscripted Pedestrian Detection

CVPR 2018poster

People detection methods are highly sensitive to occlusions between pedestrians, which are extremely frequent in many situations where cameras have to be mounted at a limited height. The reduction of camera prices allows for the generalization of static multi-camera set-ups. Using joint visual infor…

2015

Kullback-Leibler Proximal Variational Inference

NeurIPS 2015poster

We propose a new variational inference method based on the Kullback-Leibler (KL) proximal term. We make two contributions towards improving efficiency of variational inference. Firstly, we derive a KL proximal-point algorithm and show its equivalence to gradient descent with natural gradient in stoc…

Cited by 55SourcePDFScholar