← Search

Pierluigi Cassotti

6 accepted papers

2025

Towards Language-Agnostic STIPA: Universal Phonetic Transcription to Support Language Documentation at Scale

EMNLP 2025

This paper explores the use of existing state-of-the-art speech recognition models (ASR) for the task of generating narrow phonetic transcriptions using the International Phonetic Alphabet (STIPA). Unlike conventional ASR systems focused on orthographic output for high-resource languages, STIPA can

Cited by 0SourcePDFScholar
2024

Analyzing Semantic Change through Lexical Replacements

ACL 2024long

Modern language models are capable of contextualizing words based on their surrounding context. However, this capability is often compromised due to semantic change that leads to words being used in new, unexpected contexts not encountered during pre-training. In this paper, we model semantic change…

2024

More DWUGs: Extending and Evaluating Word Usage Graph Datasets in Multiple Languages

EMNLP 2024main

Word Usage Graphs (WUGs) represent human semantic proximity judgments for pairs of word uses in a weighted graph, which can be clustered to infer word sense clusters from simple pairwise word use judgments, avoiding the need for word sense definitions. SemEval-2020 Task 1 provided the first and to d…

2024

TRoTR: A Framework for Evaluating the Re-contextualization of Text Reuse

EMNLP 2024main

Current approaches for detecting text reuse do not focus on recontextualization, i.e., how the new context(s) of a reused text differs from its original context(s). In this paper, we propose a novel framework called TRoTR that relies on the notion of topic relatedness for evaluating the diachronic c…

2024

Using Synchronic Definitions and Semantic Relations to Classify Semantic Change Types

ACL 2024long

There is abundant evidence of the fact that the way words change their meaning can be classified in different types of change, highlighting the relationship between the old and new meanings (among which generalisation, specialisation and co-hyponymy transfer).In this paper, we present a way of detec…

2023

XL-LEXEME: WiC Pretrained Model for Cross-Lingual LEXical sEMantic changE

ACL 2023short

The recent introduction of large-scale datasets for the WiC (Word in Context) task enables the creation of more reliable and meaningful contextualized word embeddings.However, most of the approaches to the WiC task use cross-encoders, which prevent the possibility of deriving comparable word embeddi…