← Search

Nicolas Thome

30 accepted papers

2026

NAF: Zero-Shot Feature Upsampling via Neighborhood Attention Filtering

CVPR 2026

Vision Foundation Models (VFMs) extract spatially downsampled representations, posing challenges for pixel-level tasks. Existing upsampling approaches face a fundamental trade-off: classical filters are fast and broadly applicable but rely on fixed forms, while modern upsamplers achieve superior acc

Cited by 0SourcecodeScholar
2026

PRISM: Perception Reasoning Interleaved for Sequential Decision Making.

ICML 2026poster

Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identifies a perception–reasoning–decision gap in standalone Vision–Language Models (VLMs), which often overlook task-critical information. In this paper, we introduce …

Cited by 0SourceScholar
2025

CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation

NeurIPS 2025poster

Vision-language models (VLMs) like CLIP exhibit strong zero-shot capabilities but often fail to generalize under distribution shifts. Test-time adaptation (TTA) allows models to update at inference time without labeled data, typically via entropy minimization. However, this objective is fundamentall…

Cited by 0SourceScholar
2025

DIP: Unsupervised Dense In-Context Post-training of Visual Representations

ICCV 2025poster

We introduce DIP, a novel unsupervised post-training method designed to enhance dense representations in large-scale pretrained vision encoders for in-context scene understanding. Unlike prior approaches using complex self-distillation architectures, our method trains the vision encoder using pseudo…

2025

JAFAR: Jack up Any Feature at Any Resolution

NeurIPS 2025poster

Foundation Vision Encoders have become indispensable across a wide range of dense vision tasks. However, their operation at low spatial feature resolutions necessitates subsequent feature decompression to enable full-resolution processing. To address this limitation, we introduce JAFAR, a lightweigh…

Cited by 0SourcecodeScholar
2025

RT-HCP: Dealing with Inference Delays and Sample Efficiency to Learn Directly on Robotic Platforms

IROS 2025

Learning a controller directly on the robot requires extreme sample efficiency. Model-based reinforcement learning (RL) methods are the most sample efficient, but they often suffer from a too long inference time to meet the robot control frequency requirements. In this paper, we address the sample e

Cited by 0SourcecodeScholar
2025

Reinforcement Learning for Aligning Large Language Models Agents with Interactive Environments: Quantifying and Mitigating Prompt Overfitting

NAACL 2025findings

Reinforcement learning (RL) is a promising approach for aligning large language models (LLMs) knowledge with sequential decision-making tasks. However, few studies have thoroughly investigated the impact on LLM agents capabilities of fine-tuning them with RL in a specific environment. In this paper,…

Cited by 1SourcePDFScholar
2025

ViLU: Learning Vision-Language Uncertainties for Failure Prediction

ICCV 2025poster

Reliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new Vision-Language Uncertainty quantification framework that contextualizes uncertainty estimates by leveraging all task-relevant textual representations. V…

Cited by 0SourcePDFScholar
2024

DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut

NeurIPS 2024poster

Foundation models have emerged as powerful tools across various domains including language, vision, and multimodal tasks. While prior works have addressed unsupervised semantic segmentation, they significantly lag behind supervised models. In this paper, we use a diffusion UNet encoder as a foundati…

2024

GalLop: Learning global and local prompts for vision-language models

ECCV 2024poster

"Prompt learning has been widely adopted to efficiently adapt vision-language models (VLMs), CLIP, for few-shot image classification. Despite their success, most prompt learning methods trade-off between classification accuracy and robustness, in domain generalization or out-of-distribution (OOD) de…

2024

Supra-Laplacian Encoding for Transformer on Dynamic Graphs

NeurIPS 2024poster

Fully connected Graph Transformers (GT) have rapidly become prominent in the static graph community as an alternative to Message-Passing models, which suffer from a lack of expressivity, oversquashing, and under-reaching. However, in a dynamic context, by interconnecting all nodes at multiple snapsh…

2023

EAGLE: Large-scale Learning of Turbulent Fluid Dynamics with Mesh Transformers

ICLR 2023poster

Estimating fluid dynamics is classically done through the simulation and integration of numerical models solving the Navier-Stokes equations, which is computationally complex and time-consuming even on high-end hardware. This is a notoriously hard problem to solve, which has recently been addressed…

Cited by 36SourcePDFScholar
2023

Hybrid Energy Based Model in the Feature Space for Out-of-Distribution Detection

ICML 2023poster

Out-of-distribution (OOD) detection is a critical requirement for the deployment of deep neural networks. This paper introduces the HEAT model, a new post-hoc OOD detection method estimating the density of in-distribution (ID) samples using hybrid energy-based models (EBM) in the feature space of a…

2022

Complementing Brightness Constancy with Deep Networks for Optical Flow Prediction

ECCV 2022poster

"State-of-the-art methods for optical flow estimation rely on deep learning, which require complex sequential training schemes to reach optimal performances on real-world data. In this work, we introduce the COMBO deep network that explicitly exploits the brightness constancy (BC) model used in trad…

2022

Hierarchical Average Precision Training for Pertinent Image Retrieval

ECCV 2022poster

"Image Retrieval is commonly evaluated with Average Precision (AP) or Recall@k. Yet, those metrics, are limited to binary labels and do not take into account errors’ severity. This paper introduces a new hierarchical AP training method for pertinent image retrieval (HAPPIER). HAPPIER is based on a n…

2021

Augmenting Physical Models with Deep Networks for Complex Dynamics Forecasting

ICLR 2021oral

Forecasting complex dynamical phenomena in settings where only partial knowledge of their dynamics is available is a prevalent problem across various scientific fields. While purely data-driven approaches are arguably insufficient in this context, standard physical modeling based approaches tend to…

2021

Robust and Decomposable Average Precision for Image Retrieval

NeurIPS 2021poster

In image retrieval, standard evaluation metrics rely on score ranking, e.g. average precision (AP). In this paper, we introduce a method for robust and decomposable average precision (ROADMAP) addressing two major challenges for end-to-end training of deep neural networks with AP: non-differentiabil…

2020

Disentangling Physical Dynamics From Unknown Factors for Unsupervised Video Prediction

CVPR 2020poster

Leveraging physical knowledge described by partial differential equations (PDEs) is an appealing way to improve unsupervised video forecasting models. Since physics is too restrictive for describing the full visual content of generic video sequences, we introduce PhyDNet, a two-branch deep architect…

Cited by 404PDFcodeScholar
2020

Probabilistic Time Series Forecasting with Shape and Temporal Diversity

NeurIPS 2020poster

Probabilistic forecasting consists in predicting a distribution of possible future outcomes. In this paper, we address this problem for non-stationary time series, which is very challenging yet crucially important. We introduce the STRIPE model for representing structured diversity based on shape an…

2019

Addressing Failure Prediction by Learning Model Confidence

NeurIPS 2019poster

Assessing reliably the confidence of a deep neural net and predicting its failures is of primary importance for the practical deployment of these models. In this paper, we propose a new target criterion for model confidence, corresponding to the True Class Probability (TCP). We show how using the TC…

2019

DiscoNet: Shapes Learning on Disconnected Manifolds for 3D Editing

ICCV 2019poster

Editing 3D models is a very challenging task, as it requires complex interactions with the 3D shape to reach the targeted design, while preserving the global consistency and plausibility of the shape. In this work, we present an intelligent and user-friendly 3D editing tool, where the edited model i…

Cited by 33PDFScholar
2019

MUREL: Multimodal Relational Reasoning for Visual Question Answering

CVPR 2019poster

Multimodal attentional networks are currently state-of-the-art models for Visual Question Answering (VQA) tasks involving real images. Although attention allows to focus on the visual content relevant to the question, this simple mechanism is arguably insufficient to model complex reasoning features…

Cited by 385PDFcodeScholar
2019

Shape and Time Distortion Loss for Training Deep Time Series Forecasting Models

NeurIPS 2019poster

This paper addresses the problem of time series forecasting for non-stationary signals and multiple future steps prediction. To handle this challenging task, we introduce DILATE (DIstortion Loss including shApe and TimE), a new objective function for training deep neural networks. DILATE aims at acc…

2018

HybridNet: Classification and Reconstruction Cooperation for Semi-Supervised Learning

ECCV 2018poster

In this paper, we introduce a new model for leveraging unlabeled data to improve generalization performances of image classifiers: a two-branch encoder-decoder architecture called HybridNet. The first branch receives supervision signal and is dedicated to the extraction of invariant class-related re…

Cited by 57SourcePDFScholar
2018

Manifold Learning in Quotient Spaces

CVPR 2018poster

When learning 3D shapes we are usually interested in their intrinsic geometry rather than in their orientation. To deal with the orientation variations the usual trick consists in augmenting the data to exhibit all possible variability, and thus let the model learn both the geometry as well as the r…

Cited by 21SourcePDFScholar
2018

Revisiting Multi-Task Learning with ROCK: a Deep Residual Auxiliary Block for Visual Detection

NeurIPS 2018poster

Multi-Task Learning (MTL) is appealing for deep learning regularization. In this paper, we tackle a specific MTL context denoted as primary MTL, where the ultimate goal is to improve the performance of a given primary task by leveraging several other auxiliary tasks. Our main methodological contribu…

Cited by 69SourcePDFScholar
2017

MUTAN: Multimodal Tucker Fusion for Visual Question Answering

ICCV 2017poster

Bilinear models provide an appealing framework for mixing and merging information in Visual Question Answering (VQA) tasks. They help to learn high level associations between question meaning and visual concepts in the image, but they suffer from huge dimensionality issues. We introduce MUTAN, a mul…

Cited by 824PDFcodeScholar
2017

WILDCAT: Weakly Supervised Learning of Deep ConvNets for Image Classification, Pointwise Localization and Segmentation

CVPR 2017poster

This paper introduces WILDCAT, a deep learning method which jointly aims at aligning image regions for gaining spatial invariance and learning strongly localized features. Our model is trained using only global image labels and is devoted to three main visual recognition tasks: image classification,…

Cited by 419PDFcodeScholar
2016

WELDON: Weakly Supervised Learning of Deep Convolutional Neural Networks

CVPR 2016poster

In this paper, we introduce a novel framework for WEakly supervised Learning of Deep cOnvolutional neural Networks (WELDON). Our method is dedicated to automatically selecting relevant image regions from weak annotations, e.g. global image labels, and encompasses the following contributions. Firstly…

Cited by 217PDFcodeScholar
2015

MANTRA: Minimum Maximum Latent Structural SVM for Image Classification and Ranking

ICCV 2015poster

In this work, we propose a novel Weakly Supervised Learning (WSL) framework dedicated to learn discriminative part detectors from images annotated with a global label. Our WSL method encompasses three main contributions. Firstly, we introduce a new structured output latent variable model, Minimum m…

Cited by 43PDFScholar