2025
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees
ICML 2025poster
The propensity of large language models (LLMs) to generate hallucinations and non-factual content undermines their reliability in high-stakes domains, where rigorous control over Type I errors (the conditional probability of incorrectly classifying hallucinations as truthful content) is essential. D…