← Search

Aishwarya Ramasethu

1 accepted papers

2025

Risk Management for Mitigating Benchmark Failure Modes: BenchRisk

NeurIPS 2025poster

Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by various failure modes impacting benchmark bias, variance, coverage, or people's capacity to understand benchmark evidenc…

Cited by 0SourceScholar