← Search

Aylin Caliskan

14 accepted papers

2026

Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation

AAAI 2026technical

Large language models (LLMs) have been shown to exhibit social bias however, bias towards non-protected stigmatized identities remain understudied. Furthermore, what social features of stigmas are associated with bias in LLM outputs is unknown. From psychology literature, it has been shown that stig

Cited by 0SourcePDFScholar
2025

Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes

ACL 2025finding

To build fair AI systems we need to understand how social-group biases intrinsic to foundational encoder-based vision-language models (VLMs) manifest in biases in downstream tasks. In this study, we demonstrate that intrinsic biases in VLM representations systematically “carry over” or propagate int…

Cited by 0SourcePDFScholar
2025

Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders

NAACL 2025long

While recent work has found that vision-language models trained under the Contrastive Language Image Pre-training (CLIP) framework contain intrinsic social biases, the extent to which different upstream pre-training features of the framework relate to these biases, and hence how intrinsic bias and d…

2025

Personalized Safety in LLMs: A Benchmark and A Planning-Based Agent Approach

NeurIPS 2025poster

Large language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where user vulnerabilities differ widely. Existing safety evaluations primarily rely on context-independent metrics—such as fact…

Cited by 0SourcecodeScholar
2024

BiasDora: Exploring Hidden Biased Associations in Vision-Language Models

EMNLP 2024finding

Existing works examining Vision-Language Models (VLMs) for social biases predominantly focus on a limited set of documented bias associations, such as gender-profession or race-crime. This narrow scope often overlooks a vast range of unexamined implicit associations, restricting the identification a…

2024

Global Gallery: The Fine Art of Painting Culture Portraits through Multilingual Instruction Tuning

NAACL 2024long

Exploring the intersection of language and culture in Large Language Models (LLMs), this study critically examines their capability to encapsulate cultural nuances across diverse linguistic landscapes. Central to our investigation are three research questions: the efficacy of language-specific instr…

2024

Label-Efficient Group Robustness via Out-of-Distribution Concept Curation

CVPR 2024poster

Deep neural networks are prone to capture correlations between spurious attributes and class labels leading to low accuracy on some combinations of class labels and spurious attribute values. When a spurious attribute represents a protected class these low-accuracy groups can manifest discriminatory…

Cited by 1SourcePDFScholar
2023

'Person' == Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable Diffusion

EMNLP 2023long findings

We study stereotypes embedded within one of the most popular text-to-image generators: Stable Diffusion. We answer the question: what stereotypes of gender and nationality/continental identity does Stable Diffusion display in the absence of such information i.e. what gender and nationality/continent…

Cited by 0SourceScholar
2023

Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emotion Recognition

EMNLP 2023long findings

Previous work has established that a person's demographics and speech style affect how well speech processing models perform for them. But where does this bias come from? In this work, we present the Speech Embedding Association Test (SpEAT), a method for detecting bias in one type of model used fo…

Cited by 0SourcecodeScholar
2022

Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations

ACL 2022long

We examine the effects of contrastive visual semantic pretraining by comparing the geometry and semantic properties of contextualized English language representations formed by GPT-2 and CLIP, a zero-shot multimodal image classifier which adapts the GPT-2 architecture to encode image captions. We fi…

2022

VAST: The Valence-Assessing Semantics Test for Contextualizing Language Models

AAAI 2022technical

We introduce VAST, the Valence-Assessing Semantics Test, a novel intrinsic evaluation task for contextualized word embeddings (CWEs). Despite the widespread use of contextualizing language models (LMs), researchers have no intrinsic evaluation task for understanding the semantic quality of CWEs and…

2021

Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models

EMNLP 2021main

We use a dataset of U.S. first names with labels based on predominant gender and racial group to examine the effect of training corpus frequency on tokenization, contextualization, similarity to initial representation, and bias in BERT, GPT-2, T5, and XLNet. We show that predominantly female and non…

2021

ValNorm Quantifies Semantics to Reveal Consistent Valence Biases Across Languages and Over Centuries

EMNLP 2021main

Word embeddings learn implicit biases from linguistic regularities captured by word co-occurrence statistics. By extending methods that quantify human-like biases in word embeddings, we introduce ValNorm, a novel intrinsic evaluation task and method to quantify the valence dimension of affect in hum…