← Search

Kayo Yin

15 accepted papers

2025

SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models

NAACL 2025long

Large Language Models (LLMs) reproduce and exacerbate the social biases present in their training data, and resources to quantify this issue are limited. While research has attempted to identify and mitigate such biases, most efforts have been concentrated around English, lagging the rapid advanceme…

Cited by 1SourcePDFScholar
2024

ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles

EMNLP 2024main

Deaf and hard-of-hearing (DHH) students face significant barriers in accessing science, technology, engineering, and mathematics (STEM) education, notably due to the scarcity of STEM resources in signed languages. To help address this, we introduce ASL STEM Wiki: a parallel corpus of 254 Wikipedia a…

Cited by 1SourcePDFScholar
2024

American Sign Language Handshapes Reflect Pressures for Communicative Efficiency

ACL 2024long

Communicative efficiency is a key topic in linguistics and cognitive psychology, with many studies demonstrating how the pressure to communicate with minimal effort guides the form of natural language. However, this phenomenon is rarely explored in signed languages. This paper shows how handshapes i…

2024

Using Language Models to Disambiguate Lexical Choices in Translation

EMNLP 2024main

In translation, a concept represented by a single word in a source language can have multiple variations in a target language. The task of lexical selection requires using context to identify which variation is most appropriate for a source text. We work with native speakers of nine languages to cre…

2023

When Does Translation Require Context? A Data-driven, Multilingual Exploration

ACL 2023long

Although proper handling of discourse significantly contributes to the quality of machine translation (MT), these improvements are not adequately measured in common translation quality metrics. Recent works in context-aware MT attempt to target a small set of discourse phenomena during evaluation, h…

2021

Do Context-Aware Translation Models Pay the Right Attention?

ACL 2021long

Context-aware machine translation models are designed to leverage contextual information, but often fail to do so. As a result, they inaccurately disambiguate pronouns and polysemous words that require context for resolution. In this paper, we ask several questions: What contexts do human translator…

2021

Including Signed Languages in Natural Language Processing

ACL 2021long

Signed languages are the primary means of communication for many deaf and hard of hearing individuals. Since signed languages exhibit all the fundamental linguistic properties of natural language, we believe that tools and theories of Natural Language Processing (NLP) are crucial towards its modelin…

Cited by 130SourcePDFScholar
2021

Measuring and Increasing Context Usage in Context-Aware Machine Translation

ACL 2021long

Recent work in neural machine translation has demonstrated both the necessity and feasibility of using inter-sentential context, context from sentences other than those currently being translated. However, while many current methods present model architectures that theoretically can use this extra c…

2021

When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical Selection

EMNLP 2021main

Learning fine-grained distinctions between vocabulary items is a key challenge in learning a new language. For example, the noun “wall” has different lexical manifestations in Spanish – “pared” refers to an indoor wall while “muro” refers to an outside wall. However, this variety of lexical distinct…

Cited by 3SourcePDFScholar