← Search

Beyza Ermis

18 accepted papers

2026

Investigating Continual Pretraining in Large Language Models: Insights and Implications

ICML 2026poster

Continual learning (CL) in large language models (LLMs) is an evolving domain that focuses on developing efficient and sustainable training strategies to adapt models to emerging knowledge and achieve robustness in dynamic environments. Our primary emphasis is on continual domain-adaptive pretrainin…

Cited by 0SourceScholar
2026

Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation

ICLR 2026poster

The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage. While multilingual benchmarks have expanded, both in size and language, many rely on translations of English datasets, failing t…

Cited by 0SourcecodeScholar
2025

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

ACL 2025long

Reliable multilingual evaluation is difficult, and culturally appropriate evaluation is even harder to achieve.A common practice to fill this gap is to machine-translate English evaluation sets. However, translation introduces language bias and carries over cultural and regional assumptions from the…

Cited by 0SourcePDFScholar
2025

Multilingual Arbitration: Optimizing Data Pools to Accelerate Multilingual Progress

ACL 2025long

Synthetic data has driven recent state-of-the-art advancements, but reliance on a single oracle teacher model can lead to model collapse and bias propagation. These issues are particularly severe in multilingual settings, where no single model excels across all languages. In this study, we propose m…

Cited by 0SourcePDFScholar
2025

The State of Multilingual LLM Safety Research: From Measuring The Language Gap To Mitigating It

EMNLP 2025

This paper presents a comprehensive analysis of the linguistic diversity of LLM safety research, highlighting the English-centric nature of the field. Through a systematic review of nearly 300 publications from 2020–2024 across major NLP conferences and workshops at ACL, we identify a significant an

Cited by 0SourcePDFScholar
2024

Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?

EMNLP 2024main

In the last decade, the generalization and adaptation abilities of deep learning models were typically evaluated on fixed training and test distributions. Contrary to traditional deep learning, large language models (LLMs) are (i) even more overparameterized, (ii) trained on unlabeled text corpora c…

Cited by 1SourcePDFScholar
2024

Elo Uncovered: Robustness and Best Practices in Language Model Evaluation

NeurIPS 2024poster

In Natural Language Processing (NLP), the Elo rating system, originally designed for ranking players in dynamic games such as chess, is increasingly being used to evaluate Large Language Models (LLMs) through "A vs B" paired comparisons. However, while popular, the system's suitability for assessing…

Cited by 40SourcePDFScholar
2024

From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models

ACL 2024findings

To date, toxicity mitigation in language models has almost entirely been focused on single-language settings. As language models embrace multilingual capabilities, it’s crucial our safety measures keep pace. Recognizing this research gap, our approach expands the scope of conventional toxicity mitig…

2024

Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction Tuning

ICLR 2024poster

The Mixture of Experts (MoE) is a widely known neural architecture where an ensemble of specialized sub-models optimizes overall performance with a constant computational cost. However, conventional MoEs pose challenges at scale due to the need to store all experts in memory. In this paper, we push…

2024

The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm

EMNLP 2024main

A key concern with the concept of *“alignment”* is the implicit question of *“alignment to what?”*. AI systems are increasingly used across the world, yet safety alignment is often focused on homogeneous monolingual settings. Additionally, preference training and safety measures often overfit to har…

Cited by 18SourcePDFScholar
2023

Goodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented Models

EMNLP 2023long findings

Considerable effort has been dedicated to mitigating toxicity, but existing methods often require drastic modifications to model parameters or the use of computationally intensive auxiliary models. Furthermore, previous approaches have often neglected the crucial factor of language's evolving nature…

Cited by 0SourcecodeScholar
2023

On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research

EMNLP 2023long main

Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently retrained to address any unattended weaknesses and biases. We…

Cited by 0SourcecodeScholar
2023

PASHA: Efficient HPO and NAS with Progressive Resource Allocation

ICLR 2023poster

Hyperparameter optimization (HPO) and neural architecture search (NAS) are methods of choice to obtain the best-in-class machine learning models, but in practice they can be costly to run. When models are trained on large datasets, tuning them with HPO or NAS rapidly becomes prohibitively expensive…

2022

Memory Efficient Continual Learning with Transformers

NeurIPS 2022accept

In many real-world scenarios, data to train machine learning models becomes available over time. Unfortunately, these models struggle to continually learn new concepts without forgetting what has been learnt in the past. This phenomenon is known as catastrophic forgetting and it is difficult to prev…

Cited by 63SourcePDFScholar
2020

Linear bandits with Stochastic Delayed Feedback

ICML 2020poster

Stochastic linear bandits are a natural and well-studied model for structured exploration/exploitation problems and are widely used in applications such as on-line marketing and recommendation. One of the main challenges faced by practitioners hoping to apply existing algorithms is that usually the…

Cited by 88SourcePDFScholar
2015

Learning mixed divergences in coupled matrix and tensor factorization models

ICASSP 2015accepted

Coupled tensor factorization methods are useful for sensor fusion, combining information from several related datasets by simultaneously approximating them by products of latent tensors. In these methods, the choice of a suitable optimization criteria becomes difficult as observed datasets may have…

Cited by 0SourceScholar