← Search

Himanshu Thakur

2 accepted papers

2024

To the Cutoff... and Beyond? A Longitudinal Perspective on LLM Data Contamination

ICLR 2024poster

Recent claims about the impressive abilities of large language models (LLMs) are often supported by evaluating publicly available benchmarks. Since LLMs train on wide swaths of the internet, this practice raises concerns of data contamination, i.e., evaluating on examples that are explicitly or imp…

Cited by 27SourcePDFScholar
2023

Language Models Get a Gender Makeover: Mitigating Gender Bias with Few-Shot Data Interventions

ACL 2023short

Societal biases present in pre-trained large language models are a critical issue as these models have been shown to propagate biases in countless downstream applications, rendering them unfair towards specific groups of people. Since large-scale retraining of these models from scratch is both time…