← Search

Portia Cooper

2 accepted papers

2025

The Lies Characters Tell: Utilizing Large Language Models to Normalize Adversarial Unicode Perturbations

ACL 2025finding

Homoglyphs, Unicode characters that are visually homogeneous to Latin letters, are widely used to mask offensive content. Dynamic strategies are needed to combat homoglyphs as the Unicode library is ever-expanding and new substitution possibilities for Latin letters continuously emerge. The present…

2023

Hiding in Plain Sight: Tweets with Hate Speech Masked by Homoglyphs

EMNLP 2023short findings

To avoid detection by current NLP monitoring applications, progenitors of hate speech often replace one or more letters in offensive words with homoglyphs, visually similar Unicode characters. Harvesting real-world hate speech containing homoglyphs is challenging due to the vast replacement possibil…

Cited by 0SourceScholar