← Search

Farima Fatahi Bayat

4 accepted papers

2025

FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation

ACL 2025long

The rapid adoption of language models (LMs) across diverse applications has raised concerns about their factuality, i.e., their consistency with real-world facts. We introduce VERIFY, an evidence-based evaluation pipeline that measures LMs’ factuality in real-world user interactions. VERIFY consider…

Cited by 0SourcePDFScholar
2024

Enhanced Language Model Truthfulness with Learnable Intervention and Uncertainty Expression

ACL 2024findings

Large language models (LLMs) can generate long-form and coherent text, yet they often hallucinate facts, which undermines their reliability. To mitigate this issue, inference-time methods steer LLM representations toward the “truthful directions” previously learned for truth elicitation. However, ap…

2024

Enhancing Language Model Factuality via Activation-Based Confidence Calibration and Guided Decoding

EMNLP 2024main

Calibrating language models (LMs) aligns their generation confidence with the actual likelihood of answer correctness, which can inform users about LMs’ reliability and mitigate hallucinated content. However, prior calibration methods, such as self-consistency-based and logit-based approaches, are e…