← Search

Rodolfo Zevallos

3 accepted papers

2024

TEMA: Token Embeddings Mapping for Enriching Low-Resource Language Models

EMNLP 2024main

The objective of the research we present is to remedy the problem of the low quality of language models for low-resource languages. We introduce an algorithm, the Token Embedding Mapping Algorithm (TEMA), that maps the token embeddings of a richly pre-trained model L1 to a poorly trained model L2, t…

2023

Hints on the data for language modeling of synthetic languages with transformers

ACL 2023long

Language Models (LM) are becoming more and more useful for providing representations upon which to train Natural Language Processing applications. However, there is now clear evidence that attention-based transformers require a critical amount of language data to produce good enough LMs. The questio…

2022

WordNet-QU: Development of a Lexical Database for Quechua Varieties

COLING 2022main

In the effort to minimize the risk of extinction of a language, linguistic resources are fundamental. Quechua, a low-resource language from South America, is a language spoken by millions but, despite several efforts in the past, still lacks the resources necessary to build high-performance computat…

Cited by 4SourcePDFScholar