← Search

Iñaki San Vicente

2 accepted papers

2023

Not Enough Data to Pre-train Your Language Model? MT to the Rescue!

ACL 2023findings

In recent years, pre-trained transformer-based language models (LM) have become a key resource for implementing most NLP tasks. However, pre-training such models demands large text collections not available in most languages. In this paper, we study the use of machine-translated corpora for pre-trai…

2023

Scaling Laws for BERT in Low-Resource Settings

ACL 2023findings

Large language models are very resource intensive, both financially and environmentally, and require an amount of training data which is simply unobtainable for the majority of NLP practitioners. Previous work has researched the scaling laws of such models, but optimal ratios of model parameters, da…