2025
PHANTOM: A Benchmark for Hallucination Detection in Financial Long-Context QA
NeurIPS 2025poster
While Large Language Models (LLMs) show great promise, their tendencies to hallucinate pose significant risks in high-stakes domains like finance, especially when used for regulatory reporting and decision-making. Existing hallucination detection benchmarks fail to capture the complexities of financ…