← Search

Chrysoula Zerva

9 accepted papers

2025

A Conformal Risk Control Framework for Granular Word Assessment and Uncertainty Calibration of CLIPScore Quality Estimates

ACL 2025finding

This study explores current limitations of learned image captioning evaluation metrics, specifically the lack of granular assessments for errors within captions, and the reliance on single-point quality estimates without considering uncertainty. To address the limitations, we propose a simple yet ef…

2025

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?

NAACL 2025findings

The evaluation of image captions, looking at both linguistic fluency and semantic correspondence to visual contents, has witnessed a significant effort. Still, despite advancements such as the CLIPScore metric, multilingual captioning evaluation has remained relatively unexplored. This work presents…

2025

Rejected Dialects: Biases Against African American Language in Reward Models

NAACL 2025findings

Preference alignment via reward models helps build safe, helpful, and reliable large language models (LLMs). However, subjectivity in preference judgments and the lack of representative sampling in preference data collection can introduce new biases, hindering reward models’ fairness and equity. In…

2024

Non-Exchangeable Conformal Risk Control

ICLR 2024poster

Split conformal prediction has recently sparked great interest due to its ability to provide formally guaranteed uncertainty sets or intervals for predictions made by black-box neural models, ensuring a predefined probability of containing the actual ground truth. While the original formulation assu…

2024

”I Never Said That”: A dataset, taxonomy and baselines on response clarity classification

EMNLP 2024finding

Equivocation and ambiguity in public speech are well-studied discourse phenomena, especially in political science and analysis of political interviews. Inspired by the well-grounded theory on equivocation, we aim to resolve the closely related problem of response clarity in questions extracted from…

2023

Counterfactuals of Counterfactuals: a back-translation-inspired approach to analyse counterfactual editors

ACL 2023findings

In the wake of responsible AI, interpretability methods, which attempt to provide an explanation for the predictions of neural models have seen rapid progress. In this work, we are concerned with explanations that are applicable to natural language processing (NLP) models and tasks, and we focus spe…

2022

Disentangling Uncertainty in Machine Translation Evaluation

EMNLP 2022main

Trainable evaluation metrics for machine translation (MT) exhibit strong correlation with human judgements, but they are often hard to interpret and might produce unreliable scores under noisy or out-of-domain data. Recent work has attempted to mitigate this with simple uncertainty quantification te…

2022

Learning Disentangled Representations of Negation and Uncertainty

ACL 2022long

Negation and uncertainty modeling are long-standing tasks in natural language processing. Linguistic theory postulates that expressions of negation and uncertainty are semantically independent from each other and the content they modify. However, previous works on representation learning do not expl…

2021

Uncertainty-Aware Machine Translation Evaluation

EMNLP 2021finding

Several neural-based metrics have been recently proposed to evaluate machine translation quality. However, all of them resort to point estimates, which provide limited information at segment level. This is made worse as they are trained on noisy, biased and scarce human judgements, often resulting i…