← Search

Ishaan Watts

5 accepted papers

2026

Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting

ICML 2026poster

Standard optimizer choices for pre-training are designed to minimize pre-training loss. Yet pre-trained models are routinely subjected to further transformations—such as fine-tuning to acquire new capabilities or quantization for efficiency. In this work, we evaluate optimizer choices across model s…

Cited by 0SourceScholar
2025

RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?

AAAI 2025technical

Large language models (LLMs) and small language models (SLMs) are being adopted at remarkable speed, although their safety still remains a serious concern. With the advent of multilingual S/LLMs, the question now becomes a matter of scale: can we expand multilingual safety evaluations of these model…

2024

MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models

ACL 2024findings

Parameter efficient finetuning has emerged as a viable solution for improving the performance of Large Language Models without requiring massive resources and compute. Prior work on multilingual evaluation has shown that there is a large gap between the performance of LLMs on English and other langu…

Cited by 8SourcePDFScholar
2024

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

NAACL 2024long

There has been a surge in LLM evaluation research to understand LLM capabilities and limitations. However, much of this research has been confined to English, leaving LLM building and evaluation for non-English languages relatively unexplored. Several new LLMs have been introduced recently, necessit…

2024

PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data

EMNLP 2024main

Evaluation of multilingual Large Language Models (LLMs) is challenging due to a variety of factors – the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and the lack of local, cultural nuances in translated benchmarks. In this w…