← Search

Samuel Bell

4 accepted papers

2025

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

NeurIPS 2025poster

For Large Language Models (LLMs) to be reliably deployed in both everyday and high-stakes domains, knowing when not to answer is equally critical as answering correctly. Real-world user queries, which can be underspecified, ill-posed, or fundamentally unanswerable, require LLMs to reason about uncer…

Cited by 0SourcecodeScholar
2025

On the Role of Speech Data in Reducing Toxicity Detection Bias

NAACL 2025long

Text toxicity detection systems exhibit significant biases, producing disproportionate rates of false positives on samples mentioning demographic groups. But what about toxicity detection in speech? To investigate the extent to which text-based biases are mitigated by speech-based systems, we produc…

Cited by 0SourcePDFScholar