← Search

Riccardo Fogliato

5 accepted papers

2025

Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles

EMNLP 2025

Current benchmarks for evaluating Large Language Models (LLMs) often do not exhibit enough writing style diversity, with many adhering primarily to standardized conventions. Such benchmarks do not fully capture the rich variety of communication patterns exhibited by humans. Thus, it is possible that

Cited by 0SourcePDFScholar
2025

Stronger Neyman Regret Guarantees for Adaptive Experimental Design

ICML 2025spotlight

We study the design of adaptive, sequential experiments for unbiased average treatment effect (ATE) estimation in the design-based potential outcomes setting. Our goal is to develop adaptive designs offering *sublinear Neyman regret*, meaning their efficiency must approach that of the hindsight-opti…

2024

Multicalibration for Confidence Scoring in LLMs

ICML 2024poster

This paper proposes the use of "multicalibration": to yield interpretable and reliable confidence scores for outputs generated by large language models (LLMs). Multicalibration asks for calibration not just marginally, but simultaneously across various intersecting groupings of the data. We show how…

Cited by 18SourcePDFScholar
2024

Precise Model Benchmarking with Only a Few Observations

EMNLP 2024main

How can we precisely estimate a large language model’s (LLM) accuracy on questions belonging to a specific topic within a larger question-answering dataset? The standard direct estimator, which averages the model’s accuracy on the questions in each subgroup, may exhibit high variance for subgroups (…

Cited by 1SourcePDFScholar
2020

Fairness Evaluation in Presence of Biased Noisy Labels

AISTATS 2020poster

Risk assessment tools are widely used around the country to inform decision making within the criminal justice system. Recently, considerable attention has been devoted to the question of whether such tools may suffer from racial bias. In this type of assessment, a fundamental issue is that the trai…