← Search

Aryan Shrivastava

2 accepted papers

2026

Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare

ICLR 2026poster

Current medical language model (LM) benchmarks often over-simplify the complexities of day-to-day clinical practice tasks and instead rely on evaluating LMs on multiple-choice board exam questions. In psychiatry especially, these challenges are worsened by fairness and bias issues, since models can…

Cited by 0SourcecodeScholar
2025

Absence Bench: Language Models Can’t See What’s Missing

NeurIPS 2025spotlight

Large language models (LLMs) are increasingly capable of processing long inputs and locating specific information within them, as evidenced by their performance on the Needle in a Haystack (NIAH) test. However, while models excel at recalling surprising information, they still struggle to identify c…

Cited by 0SourceScholar