← Search

Tatsuya Aoyama

8 accepted papers

2026

Predicting the Emergence of Induction Heads in Language Model Pretraining

ICML 2026poster

Specialized attention heads dubbed induction heads (IHs) have been argued to underlie the remarkable in-context learning capabilities of modern language models; yet, a precise characterization of their emergence, especially in the context of language modeling, remains wanting. In this study, we inve…

Cited by 0SourceScholar
2025

Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs

ACL 2025long

Do language models (LMs) offer insights into human language learning? A common argument against this idea is that because their architecture and training paradigm are so vastly different from humans, LMs can learn arbitrary inputs as easily as natural languages. We test this claim by training LMs to…

Cited by 0SourcePDFScholar
2025

Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning

EMNLP 2025

Humans have a remarkable ability to acquire and understand grammatical phenomena that are seen rarely, if ever, during childhood. Recent evidence suggests that language models with human-scale pretraining data may possess a similar ability by generalizing from frequent to rare constructions. However

Cited by 0SourcePDFScholar
2024

DISRPT: A Multilingual, Multi-domain, Cross-framework Benchmark for Discourse Processing

COLING 2024main

This paper presents DISRPT, a multilingual, multi-domain, and cross-framework benchmark dataset for discourse processing, covering the tasks of discourse unit segmentation, connective identification, and relation classification. DISRPT includes 13 languages, with data from 24 corpora covering about…

2024

GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains

EMNLP 2024main

Work on shallow discourse parsing in English has focused on the Wall Street Journal corpus, the only large-scale dataset for the language in the PDTB framework. However, the data is not openly available, is restricted to the news domain, and is by now 35 years old. In this paper, we present and eval…

2024

J-SNACS: Adposition and Case Supersenses for Japanese Joshi

COLING 2024main

Many languages use adpositions (prepositions or postpositions) to mark a variety of semantic relations, with different languages exhibiting both commonalities and idiosyncrasies in the relations grouped under the same lexeme. We present the first Japanese extension of the SNACS framework (Schneider…