← Search

Silvia Severini

4 accepted papers

2024

SilverAlign: MT-Based Silver Data Algorithm for Evaluating Word Alignment

COLING 2024main

Word alignments are essential for a variety of NLP tasks. Therefore, choosing the best approaches for their creation is crucial. However, the scarce availability of gold evaluation data makes the choice difficult. We propose SilverAlign, a new method to automatically create silver data for the evalu…

2023

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

ACL 2023long

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., making them better for about 100 languages. We instead scale LLMs horizontally: we create, through continued pretraining, Glot500-m, an LLM that covers 511 predominantly low-resource languages. An importan…

2022

Graph-Based Multilingual Label Propagation for Low-Resource Part-of-Speech Tagging

EMNLP 2022main

Part-of-Speech (POS) tagging is an important component of the NLP pipeline, but many low-resource languages lack labeled data for training. An established method for training a POS tagger in such a scenario is to create a labeled training set by transferring from high-resource languages. In this pap…

2020

Combining Word Embeddings with Bilingual Orthography Embeddings for Bilingual Dictionary Induction

COLING 2020main

Bilingual dictionary induction (BDI) is the task of accurately translating words to the target language. It is of great importance in many low-resource scenarios where cross-lingual training data is not available. To perform BDI, bilingual word embeddings (BWEs) are often used due to their low bilin…

Cited by 6SourcePDFScholar