← Search

Sebastin Santy

4 accepted papers

2025

Position: When Incentives Backfire, Data Stops Being Human

ICML 2025poster

Progress in AI has relied on human-generated data, from annotator marketplaces to the wider Internet. However, the widespread use of large language models now threatens the quality and integrity of human-generated data on these very platforms. We argue that this issue goes beyond the immediate chall…

Cited by 0SourcePDFScholar
2025

Semantic and Expressive Variations in Image Captions Across Languages

CVPR 2025poster

Most vision-language models today are primarily trained on English image-text pairs, with non-English pairs often filtered out. Evidence from cross-cultural psychology suggests that this approach will bias models against perceptual modes exhibited by people who speak other (non-English) languages. W…

Cited by 0SourcePDFScholar
2024

Multilingual Diversity Improves Vision-Language Representations

NeurIPS 2024spotlight

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however, have been shown to be English-centric (e.g., ImageNet). Cons…

Cited by 8SourcePDFScholar
2023

NLPositionality: Characterizing Design Biases of Datasets and Models

ACL 2023long

Design biases in NLP systems, such as performance differences for different populations, often stem from their creator’s positionality, i.e., views and lived experiences shaped by identity and background. Despite the prevalence and risks of design biases, they are hard to quantify because researcher…