← Search

Sshubam Verma

2 accepted papers

2025

MILU: A Multi-task Indic Language Understanding Benchmark

NAACL 2025long

Evaluating Large Language Models (LLMs) in low-resource and linguistically diverse languages remains a significant challenge in NLP, particularly for languages using non-Latin scripts like those spoken in India. Existing benchmarks predominantly focus on English, leaving substantial gaps in assessin…

2024

Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

EMNLP 2024main

Large Language Models (LLMs) are increasingly relied upon to evaluate text outputs of other LLMs, thereby influencing leaderboards and development decisions. However, concerns persist over the accuracy of these assessments and the potential for misleading conclusions. In this work, we investigate th…