← Search

Svetlana Kiritchenko

8 accepted papers

2025

Tackling Social Bias against the Poor: a Dataset and a Taxonomy on Aporophobia

NAACL 2025findings

Eradicating poverty is the first goal in the U.N. Sustainable Development Goals. However, aporophobia – the societal bias against people living in poverty – constitutes a major obstacle to designing, approving and implementing poverty-mitigation policies. This work presents an initial step towards o…

Cited by 0SourcePDFScholar
2025

Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals

NAACL 2025long

With the advent of Large Language Models (LLMs) possessing increasingly impressive capabilities, a number of Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs. Such models condition generated text on both an input image and a text prompt, enabling a variety o…

Cited by 8SourcePDFScholar
2025

When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Text

ACL 2025finding

Detecting AI-generated text is a difficult problem to begin with; detecting AI-generated text on social media is made even more difficult due to the short text length and informal, idiosyncratic language of the internet. It is nonetheless important to tackle this problem, as social media represents…

2024

Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse

EMNLP 2024main

This work provides an explanatory view of how LLMs can apply moral reasoning to both criticize and defend sexist language. We assessed eight large language models, all of which demonstrated the capability to provide explanations grounded in varying moral perspectives for both critiquing and endorsin…

2024

Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes

COLING 2024main

Gender stereotypes are pervasive beliefs about individuals based on their gender that play a significant role in shaping societal attitudes, behaviours, and even opportunities. Recognizing the negative implications of gender stereotypes, particularly in online communications, this study investigates…

Cited by 3SourcePDFScholar
2022

Improving Generalizability in Implicitly Abusive Language Detection with Concept Activation Vectors

ACL 2022long

Robustness of machine learning models on ever-changing real-world data is critical, especially for applications affecting human well-being such as content moderation. New kinds of abusive language continually emerge in online discussions in response to current events (e.g., COVID-19), and the deploy…

2022

Necessity and Sufficiency for Explaining Text Classifiers: A Case Study in Hate Speech Detection

NAACL 2022long

We present a novel feature attribution method for explaining text classifiers, and analyze it in the context of hate speech detection. Although feature attribution models usually provide a single importance score for each token, we instead provide two complementary and theoretically-grounded scores…

2021

Understanding and Countering Stereotypes: A Computational Approach to the Stereotype Content Model

ACL 2021long

Stereotypical language expresses widely-held beliefs about different social categories. Many stereotypes are overtly negative, while others may appear positive on the surface, but still lead to negative consequences. In this work, we present a computational approach to interpreting stereotypes in te…

Cited by 44SourcePDFScholar