← Search

CELINE HUDELOT

8 accepted papers

2026

Should We Still Pretrain Encoders with Masked Language Modeling?

ICLR 2026poster

Learning high-quality text representations is fundamental to a wide range of NLP tasks. While encoder pretraining has traditionally relied on Masked Language Modeling (MLM), recent evidence suggests that decoder models pretrained with Causal Language Modeling (CLM) can be effectively repurposed as e…

Cited by 0SourceScholar
2025

ColPali: Efficient Document Retrieval with Vision Language Models

ICLR 2025poster

Documents are visually rich structures that convey information through text, but also figures, page layouts, tables, or even fonts. Since modern retrieval systems mainly rely on the textual information they extract from document pages to index documents -often through lengthy and brittle processes-,…

Cited by 43SourcePDFScholar
2025

Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings

EMNLP 2025

A limitation of modern document retrieval embedding methods is that they typically encode passages (chunks) from the same documents independently, often overlooking crucial contextual information from the rest of the document that could greatly improve individual chunk representations.In this work,

2023

Open-Set Likelihood Maximization for Few-Shot Learning

CVPR 2023poster

We tackle the Few-Shot Open-Set Recognition (FSOSR) problem, i.e. classifying instances among a set of classes for which we only have a few labeled samples, while simultaneously detecting instances that do not belong to any known class. We explore the popular transductive setting, which leverages th…

2023

Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial Applications

EMNLP 2023short main

Instruction Fine-Tuning (IFT) is a powerful paradigm that strengthens the zero-shot capabilities of Large Language Models (LLMs), but in doing so induces new evaluation metric requirements. We show LLM-based metrics to be well adapted to these requirements, and leverage them to conduct an investigat…

Cited by 0SourcecodeScholar
2022

Towards Job-Transition-Tag Graph for a Better Job Title Representation Learning

NAACL 2022findings

Works on learning job title representation are mainly based on Job-Transition Graph, built from the working history of talents. However, since these records are usually messy, this graph is very sparse, which affects the quality of the learned representation and hinders further analysis. To address…

2020

Semi-Supervised Semantic Segmentation With Cross-Consistency Training

CVPR 2020poster

In this paper, we present a novel cross-consistency based semi-supervised approach for semantic segmentation. Consistency training has proven to be a powerful semi-supervised learning framework for leveraging unlabeled data under the cluster assumption, in which the decision boundary should lie in l…

Cited by 1026PDFcodeScholar
2017

MuCaLe-Net: Multi Categorical-Level Networks to Generate More Discriminating Features

CVPR 2017poster

In a transfer-learning scheme, the intermediate layers of a pre-trained CNN are employed as universal image representation to tackle many visual classification problems. The current trend to generate such representation is to learn a CNN on a large set of images labeled among the most specific categ…

Cited by 25PDFScholar