← Search

Nathan Noiry

7 accepted papers

2024

Towards More Robust NLP System Evaluation: Handling Missing Scores in Benchmarks

EMNLP 2024finding

The evaluation of natural language processing (NLP) systems is crucial for advancing the field, but current benchmarking approaches often assume that all systems have scores available for all tasks, which is not always practical. In reality, several factors such as the cost of running baseline, priv…

Cited by 5SourcePDFScholar
2022

Beyond Mahalanobis Distance for Textual OOD Detection

NeurIPS 2022accept

As the number of AI systems keeps growing, it is fundamental to implement and develop efficient control mechanisms to ensure the safe and proper functioning of machine learning (ML) systems. Reliable out-of-distribution (OOD) detection aims to detect test samples that are statistically far from the…

Cited by 53SourcePDFScholar
2022

Learning Disentangled Textual Representations via Statistical Measures of Similarity

ACL 2022long

When working with textual data, a natural application of disentangled representations is the fair classification where the goal is to make predictions without being biased (or influenced) by sensible attributes that may be present in the data (e.g., age, gender or race). Dominant approaches to disen…

2022

Mitigating Gender Bias in Face Recognition using the von Mises-Fisher Mixture Model

ICML 2022spotlight

In spite of the high performance and reliability of deep learning algorithms in a wide range of everyday applications, many investigations tend to show that a lot of models exhibit biases, discriminating against specific subgroups of the population (e.g. gender, ethnicity). This urges the practition…

2022

What are the best Systems? New Perspectives on NLP Benchmarking

NeurIPS 2022accept

In Machine Learning, a benchmark refers to an ensemble of datasets associated with one or multiple metrics together with a way to aggregate different systems performances. They are instrumental in {\it (i)} assessing the progress of new methods along different axes and {\it (ii)} selecting the best…

2021

Learning from Biased Data: A Semi-Parametric Approach

ICML 2021spotlight

We consider risk minimization problems where the (source) distribution $P_S$ of the training observations $Z_1, \ldots, Z_n$ differs from the (target) distribution $P_T$ involved in the risk that one seeks to minimize. Under the natural assumption that $P_S$ dominates $P_T$, \textit{i.e.} $P_T< \! \…

Cited by 11SourcePDFScholar
2021

Online Matching in Sparse Random Graphs: Non-Asymptotic Performances of Greedy Algorithm

NeurIPS 2021poster

Motivated by sequential budgeted allocation problems, we investigate online matching problems where connections between vertices are not i.i.d., but they have fixed degree distributions -- the so-called configuration model. We estimate the competitive ratio of the simplest algorithm, GREEDY, by app…

Cited by 6SourcePDFScholar