← Search

Edgar Altszyler

3 accepted papers

2025

MessIRve: A Large-Scale Spanish Information Retrieval Dataset

EMNLP 2025

Information retrieval (IR) is the task of finding relevant documents in response to a user query. Although Spanish is the second most spoken native language, there are few Spanish IR datasets, which limits the development of information access tools for Spanish speakers. We introduce MessIRve, a lar

2023

On the Interpretability and Significance of Bias Metrics in Texts: a PMI-based Approach

ACL 2023short

In recent years, word embeddings have been widely used to measure biases in texts. Even if they have proven to be effective in detecting a wide variety of biases, metrics based on word embeddings lack transparency and interpretability. We analyze an alternative PMI-based metric to quantify biases in…

2022

The Undesirable Dependence on Frequency of Gender Bias Metrics Based on Word Embeddings

EMNLP 2022finding

Numerous works use word embedding-based metrics to quantify societal biases and stereotypes in texts. Recent studies have found that word embeddings can capture semantic similarity but may be affected by word frequency. In this work we study the effect of frequency when measuring female vs. male gen…