← Search

Flor Miriam Plaza-del-Arco

10 accepted papers

2025

La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America

ACL 2025long

Leaderboards showcase the current capabilities and limitations of Large Language Models (LLMs). To motivate the development of LLMs that represent the linguistic and cultural diversity of the Spanish-speaking community, we present La Leaderboard, the first open-source leaderboard to evaluate generat…

Cited by 0SourcePDFScholar
2025

Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks

NAACL 2025long

As Large Language Models (LLMs) continue to evolve, evaluating them remains a persistent challenge. Many recent evaluations use LLMs as judges to score outputs from other LLMs, often relying on a single large model like GPT-4o. However, using a single LLM judge is prone to intra-model bias, and many…

2025

MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation

EMNLP 2025

Ensuring the moral reasoning capabilities of Large Language Models (LLMs) is a growing concern as these systems are used in socially sensitive tasks. Nevertheless, current evaluation benchmarks present two major shortcomings: a lack of annotations that justify moral classifications, which limits tra

2025

Seeing Race, Feeling Bias: Emotion Stereotyping in Multimodal Language Models

EMNLP 2025

Large language models (LLMs) are increasingly used to predict human emotions, but previous studies show that these models reproduce gendered emotion stereotypes. Emotion stereotypes are also tightly tied to race and skin tone (consider for example the trope of the angry black woman), but previous wo

Cited by 4SourcePDFScholar
2025

Your Mileage May Vary: How Empathy and Demographics Shape Human Preferences in LLM Responses

EMNLP 2025

As large language models (LLMs) increasingly assist in subjective decision-making (e.g., moral reasoning, advice), it is critical to understand whose preferences they align with—and why. While prior work uses aggregate human judgments, demographic variation and its linguistic drivers remain underexp

Cited by 0SourcePDFScholar
2024

Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution

ACL 2024long

Large language models (LLMs) reflect societal norms and biases, especially about gender. While societal biases and stereotypes have been extensively researched in various NLP applications, there is a surprising gap for emotion analysis. However, emotion and gender are closely linked in societal disc…

2024

Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models

EMNLP 2024finding

Emotions play important epistemological and cognitive roles in our lives, revealing our values and guiding our actions. Previous work has shown that LLMs display biases in emotion attribution along gender lines. However, unlike gender, which says little about our values, religion, as a socio-cultura…

Cited by 6SourcePDFScholar
2024

Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions

COLING 2024main

Emotions are a central aspect of communication. Consequently, emotion analysis (EA) is a rapidly growing field in natural language processing (NLP). However, there is no consensus on scope, direction, or methods. In this paper, we conduct a thorough review of 154 relevant NLP publications from the l…

2024

MentalRiskES: A New Corpus for Early Detection of Mental Disorders in Spanish

COLING 2024main

With mental health issues on the rise on the Web, especially among young people, there is a growing need for effective identification and intervention. In this paper, we introduce a new open-sourced corpus for the early detection of mental disorders in Spanish, focusing on eating disorders, depressi…

2022

Natural Language Inference Prompts for Zero-shot Emotion Classification in Text across Corpora

COLING 2022main

Within textual emotion classification, the set of relevant labels depends on the domain and application scenario and might not be known at the time of model development. This conflicts with the classical paradigm of supervised learning in which the labels need to be predefined. A solution to obtain…