← Search

Sophie Hao

6 accepted papers

2025

ERAS: Evaluating the Robustness of Chinese NLP Models to Morphological Garden Path Errors

NAACL 2025long

In languages without orthographic word boundaries, NLP models perform _word segmentation_, either as an explicit preprocessing step or as an implicit step in an end-to-end computation. This paper shows that Chinese NLP models are vulnerable to _morphological garden path errors_—errors caused by a fa…

2025

ModelCitizens: Representing Community Voices in Online Safety

EMNLP 2025

Automatic toxic language detection is important for creating safe, inclusive online spaces. However, it is a highly subjective task, with perceptions of toxic language shaped by community norms and lived experience. Existing toxicity detection models are typically trained on annotations that collaps

Cited by 0SourcePDFScholar
2025

What Goes Into a LM Acceptability Judgment? Rethinking the Impact of Frequency and Length

NAACL 2025long

When comparing the linguistic capabilities of language models (LMs) with humans using LM probabilities, factors such as the length of the sequence and the unigram frequency of lexical items have a significant effect on LM probabilities in ways that humans are largely robust to. Prior works in compar…

2024

Reflecting the Male Gaze: Quantifying Female Objectification in 19th and 20th Century Novels

COLING 2024main

Inspired by the concept of the male gaze (Mulvey, 1975) in literature and media studies, this paper proposes a framework for analyzing gender bias in terms of female objectification—the extent to which a text portrays female individuals as objects of visual pleasure. Our framework measures female ob…

2023

Verb Conjugation in Transformers Is Determined by Linear Encodings of Subject Number

EMNLP 2023short findings

Deep architectures such as Transformers are sometimes criticized for having uninterpretable "black-box" representations. We use causal intervention analysis to show that, in fact, some linguistic features are represented in a linear, interpretable format. Specifically, we show that BERT's ability to…

Cited by 0SourcecodeScholar