← Search

Robert Morabito

4 accepted papers

2025

Fine-Tuned LLMs are “Time Capsules” for Tracking Societal Bias Through Books

NAACL 2025long

Books, while often rich in cultural insights, can also mirror societal biases of their eras—biases that Large Language Models (LLMs) may learn and perpetuate during training. We introduce a novel method to trace and quantify these biases using fine-tuned LLMs. We develop BookPAGE, a corpus comprisin…

2024

Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models

ACL 2024long

As the use of Large Language Models (LLMs) becomes more widespread, understanding their self-evaluation of confidence in generated responses becomes increasingly important as it is integral to the reliability of the output of these models. We introduce the concept of Confidence-Probability Alignment…

2024

STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions

EMNLP 2024main

Mitigating explicit and implicit biases in Large Language Models (LLMs) has become a critical focus in the field of natural language processing. However, many current methodologies evaluate scenarios in isolation, without considering the broader context or the spectrum of potential biases within eac…

2023

Debiasing should be Good and Bad: Measuring the Consistency of Debiasing Techniques in Language Models

ACL 2023findings

Debiasing methods that seek to mitigate the tendency of Language Models (LMs) to occasionally output toxic or inappropriate text have recently gained traction. In this paper, we propose a standardized protocol which distinguishes methods that yield not only desirable results, but are also consistent…