← Search

Lorenzo Pacchiardi

3 accepted papers

2025

Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture

IJCAI 2025

Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation, adopting conflicting terminologies, and overlooking each other's co

Cited by 0SourcePDFScholar
2025

PredictaBoard: Benchmarking LLM Score Predictability

ACL 2025finding

Despite possessing impressive skills, Large Language Models (LLMs) often fail unpre-dictably, demonstrating inconsistent success in even basic common sense reasoning tasks. This unpredictability poses a significant challenge to ensuring their safe deployment, as identifying and operating within a re…

2024

How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions

ICLR 2024poster

Large language models (LLMs) can “lie”, which we define as outputting false statements when incentivised to, despite “knowing” the truth in a demonstrable sense. LLMs might “lie”, for example, when instructed to output misinformation. Here, we develop a simple lie detector that requires neither acce…