← Search

Sunayana Sitaram

21 accepted papers

2025

A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications

EMNLP 2025

Misgendering is the act of referring to someone by a gender that does not match their chosen identity. It marginalizes and undermines a person’s sense of self, causing significant harm. English-based approaches have clear-cut approaches to avoiding misgendering, such as the use of the pronoun “they”

2025

Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs

COLING 2025main

Large language models (LLMs) have revolutionized various domains but still struggle with non-Latin scripts and low-resource languages. This paper addresses the critical challenge of improving multilingual performance without extensive fine-tuning. We introduce a novel dynamic learning approach that…

Cited by 0SourcePDFScholar
2025

Improving Consistency in LLM Inference using Probabilistic Tokenization

NAACL 2025findings

Prior research has demonstrated noticeable performance gains through the use of probabilistic tokenizations, an approach that involves employing multiple tokenizations of the same input string during the training phase of a language model. Despite these promising findings, modern large language mode…

Cited by 0SourcePDFScholar
2025

Improving Cross Lingual Transfer by Pretraining with Active Forgetting

EMNLP 2025

Large Language Models (LLMs) demonstrate exceptional capabilities in a multitude of NLP tasks. However, the efficacy of such models to languages other than English is often limited. Prior works have shown that encoder-only models such as BERT or XLM-RoBERTa show impressive cross lingual transfer of

Cited by 0SourcePDFScholar
2024

A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models

EMNLP 2024finding

Vision-language models (VLMs) have gained widespread adoption in both industry and academia. In this study, we propose a unified framework for systematically evaluating gender, race, and age biases in VLMs with respect to professions. Our evaluation encompasses all supported inference modes of the r…

Cited by 10SourcePDFScholar
2024

Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting

EMNLP 2024main

Socio-demographic prompting is a commonly employed approach to study cultural biases in LLMs as well as for aligning models to certain cultures. In this paper, we systematically probe four LLMs (Llama 3, Mistral v0.2, GPT-3.5 Turbo and GPT4) with prompts that are conditioned on culturally sensitive…

Cited by 7SourcePDFScholar
2024

CultureLLM: Incorporating Cultural Differences into Large Language Models

NeurIPS 2024poster

Large language models (LLMs) have been observed to exhibit bias towards certain cultures due to the predominance of training data obtained from English corpora. Considering that multilingual cultural data is often expensive to procure, existing methodologies address this challenge through prompt eng…

2024

DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures

COLING 2024main

Generative models are increasingly being used in various applications, such as text generation, commonsense reasoning, and question-answering. To be effective globally, these models must be aware of and account for local socio-cultural contexts, making it necessary to have benchmarks to evaluate the…

Cited by 10SourcePDFScholar
2024

M5 – A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks

EMNLP 2024finding

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). Despite their impressive capabilities, LLMs often exhibit significant performance…

2024

MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models

ACL 2024findings

Parameter efficient finetuning has emerged as a viable solution for improving the performance of Large Language Models without requiring massive resources and compute. Prior work on multilingual evaluation has shown that there is a large gap between the performance of LLMs on English and other langu…

Cited by 8SourcePDFScholar
2024

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

NAACL 2024long

There has been a surge in LLM evaluation research to understand LLM capabilities and limitations. However, much of this research has been confined to English, leaving LLM building and evaluation for non-English languages relatively unexplored. Several new LLMs have been introduced recently, necessit…

2024

METAL: Towards Multilingual Meta-Evaluation

NAACL 2024findings

With the rising human-like precision of Large Language Models (LLMs) in numerous tasks, their utilization in a variety of real-world applications is becoming more prevalent. Several studies have shown that LLMs excel on many standard NLP benchmarks. However, it is challenging to evaluate LLMs due to…

2024

PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data

EMNLP 2024main

Evaluation of multilingual Large Language Models (LLMs) is challenging due to a variety of factors – the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and the lack of local, cultural nuances in translated benchmarks. In this w…

2024

Teaching LLMs to Abstain across Languages via Multilingual Feedback

EMNLP 2024main

Multilingual LLMs often have knowledge disparities across languages, with larger gaps in under-resourced languages. Teaching LLMs to abstain in the face of knowledge gaps is thus a promising strategy to mitigate hallucinations in multilingual settings. However, previous studies on LLM abstention pri…

2023

A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models

ACL 2023long

Compression techniques for deep learning have become increasingly popular, particularly in settings where latency and memory constraints are imposed. Several methods, such as pruning, distillation, and quantization, have been adopted for compressing models, each providing distinct advantages. Howeve…

2023

Analysing the Masked Predictive Coding Training Criterion for Pre-Training a Speech Representation Model

ICASSP 2023accepted

Recent developments in pre-trained speech representation utilizing self-supervised learning (SSL) have yielded exceptional results on a variety of downstream tasks. One such technique, known as masked predictive coding (MPC), has been employed by some of the most high-performing models. In this stud…

Cited by 0SourceScholar
2023

MEGA: Multilingual Evaluation of Generative AI

EMNLP 2023long main

Generative AI models have shown impressive performance on many Natural Language Processing tasks such as language understanding, reasoning, and language generation. An important question being asked by the AI community today is about the capabilities and limits of these models, and it is clear that…

Cited by 0SourceScholar
2023

On Evaluating and Mitigating Gender Biases in Multilingual Settings

ACL 2023findings

While understanding and removing gender biases in language models has been a long-standing problem in Natural Language Processing, prior research work has primarily been limited to English. In this work, we investigate some of the challenges with evaluating and mitigating biases in multilingual sett…

Cited by 21SourcePDFScholar
2023

Representativeness as a Forgotten Lesson for Multilingual and Code-switched Data Collection and Preparation

EMNLP 2023long findings

Multilingualism is widespread around the world and code-switching (CSW) is a common practice among different language pairs/tuples across locations and regions. However, there is still not much progress in building successful CSW systems, despite the recent advances in Massive Multilingual Language…

Cited by 0SourceScholar
2022

On the Calibration of Massively Multilingual Language Models

EMNLP 2022main

Massively Multilingual Language Models (MMLMs) have recently gained popularity due to their surprising effectiveness in cross-lingual transfer. While there has been much work in evaluating these models for their performance on a variety of tasks and languages, little attention has been paid on how w…

2021

A Survey of Code-switching: Linguistic and Social Perspectives for Language Technologies

ACL 2021long

The analysis of data in which multiple languages are represented has gained popularity among computational linguists in recent years. So far, much of this research focuses mainly on the improvement of computational methods and largely ignores linguistic and social aspects of C-S discussed across a w…

Cited by 78SourcePDFScholar