← Search

Varun Gumma

5 accepted papers

2026

OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!

ICLR 2026poster

Large Language Model (LLM) safety is one of the most pressing challenges for enabling wide-scale deployment. While most studies and global discussions focus on generic harms, such as models assisting users in harming themselves or others, enterprises face a more fundamental concern: whether LLM-base…

Cited by 0SourcecodeScholar
2025

Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models

NAACL 2025long

Neural Machine Translation (NMT) models have traditionally used Sinusoidal Positional Embeddings (PEs), which often struggle to capture long-range dependencies and are inefficient for handling extended context or document-level translation tasks. This work addresses the challenge of transitioning pr…

2024

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

NAACL 2024long

There has been a surge in LLM evaluation research to understand LLM capabilities and limitations. However, much of this research has been confined to English, leaving LLM building and evaluation for non-English languages relatively unexplored. Several new LLMs have been introduced recently, necessit…

2024

METAL: Towards Multilingual Meta-Evaluation

NAACL 2024findings

With the rising human-like precision of Large Language Models (LLMs) in numerous tasks, their utilization in a variety of real-world applications is becoming more prevalent. Several studies have shown that LLMs excel on many standard NLP benchmarks. However, it is challenging to evaluate LLMs due to…

2024

PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data

EMNLP 2024main

Evaluation of multilingual Large Language Models (LLMs) is challenging due to a variety of factors – the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and the lack of local, cultural nuances in translated benchmarks. In this w…