← Search

Ryo Nagata

8 accepted papers

2025

A New Formulation of Zipf’s Meaning-Frequency Law through Contextual Diversity

ACL 2025long

This paper proposes formulating Zipf’s meaning-frequency law, the power law between word frequency and the number of meanings, as a relationship between word frequency and contextual diversity. The proposed formulation quantifies meaning counts as contextual diversity, which is based on the directio…

2025

Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport

ACL 2025long

Lexical semantic change detection aims to identify shifts in word meanings over time. While existing methods using embeddings from a diachronic corpus pair estimate the degree of change for target words, they offer limited insight into changes at the level of individual usage instances. To address t…

2024

A Computational Approach to Quantifying Grammaticization of English Deverbal Prepositions

COLING 2024main

This paper explores grammaticization of deverbal prepositions by a computational approach based on corpus data. Deverbal prepositions are words or phrases that are derived from a verb and that behave as a preposition such as “regarding” and “according to”. Linguistic studies have revealed important…

Cited by 0SourcePDFScholar
2023

Variance Matters: Detecting Semantic Differences without Corpus/Word Alignment

EMNLP 2023long main

In this paper, we propose methods for discovering semantic differences in words appearing in two corpora. The key idea is to measure the coverage of meanings of a word in a corpus through the norm of its mean word vector, which is equivalent to examining a kind of variance of the word vector distrib…

Cited by 0SourceScholar
2022

Exploring the Capacity of a Large-scale Masked Language Model to Recognize Grammatical Errors

ACL 2022findings

In this paper, we explore the capacity of a language model-based method for grammatical error detection in detail. We first show that 5 to 10% of training data are enough for a BERT-based error detection method to achieve performance equivalent to what a non-language model-based method can achieve w…

Cited by 7SourcePDFScholar
2022

Revisiting Statistical Laws of Semantic Shift in Romance Cognates

COLING 2022main

This article revisits statistical relationships across Romance cognates between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy. Cognates are words that are derived from a common etymon, in this case, a Latin ancestor. Despite their shared etymology, some co…

Cited by 4SourcePDFScholar
2021

Exploring Methods for Generating Feedback Comments for Writing Learning

EMNLP 2021main

The task of generating explanatory notes for language learners is known as feedback comment generation. Although various generation techniques are available, little is known about which methods are appropriate for this task. Nagata (2019) demonstrates the effectiveness of neural-retrieval-based meth…

2020

Taking the Correction Difficulty into Account in Grammatical Error Correction Evaluation

COLING 2020main

This paper presents performance measures for grammatical error correction which take into account the difficulty of error correction. To the best of our knowledge, no conventional measure has such functionality despite the fact that some errors are easy to correct and others are not. The main purpos…