← Search

Aitor Ormazabal

6 accepted papers

2025

Improving the Efficiency of Visually Augmented Language Models

COLING 2025main

Despite the impressive performance of autoregressive Language Models (LM) it has been shown that due to reporting bias, LMs lack visual knowledge, i.e. they do not know much about the visual world and its properties. To augment LMs with visual knowledge, existing solutions often rely on explicit ima…

2024

Latxa: An Open Language Model and Evaluation Suite for Basque

ACL 2024long

We introduce Latxa, a family of large language models for Basque ranging from 7 to 70 billion parameters. Latxa is based on Llama 2, which we continue pretraining on a new Basque corpus comprising 4.3M documents and 4.2B tokens. Addressing the scarcity of high-quality benchmarks for Basque, we furth…

2023

CombLM: Adapting Black-Box Language Models through Small Fine-Tuned Models

EMNLP 2023long main

Methods for adapting language models (LMs) to new tasks and domains have traditionally assumed white-box access to the model, and work by modifying its parameters. However, this is incompatible with a recent trend in the field, where the highest quality models are only available as black-boxes throu…

Cited by 0SourceScholar
2022

PoeLM: A Meter- and Rhyme-Controllable Language Model for Unsupervised Poetry Generation

EMNLP 2022finding

Formal verse poetry imposes strict constraints on the meter and rhyme scheme of poems. Most prior work on generating this type of poetry uses existing poems for supervision, which are difficult to obtain for most languages and poetic forms. In this work, we propose an unsupervised approach to genera…

2022

Principled Paraphrase Generation with Parallel Corpora

ACL 2022long

Round-trip Machine Translation (MT) is a popular choice for paraphrase generation, which leverages readily available parallel corpora for supervision. In this paper, we formalize the implicit similarity function induced by this approach, and show that it is susceptible to non-paraphrase pairs sharin…

2021

Beyond Offline Mapping: Learning Cross-lingual Word Embeddings through Context Anchoring

ACL 2021long

Recent research on cross-lingual word embeddings has been dominated by unsupervised mapping approaches that align monolingual embeddings. Such methods critically rely on those embeddings having a similar structure, but it was recently shown that the separate training in different languages causes de…

Cited by 15SourcePDFScholar