← Search

Elizaveta Korotkova

2 accepted papers

2024

BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training

EMNLP 2024main

Language models can greatly benefit from efficient tokenization. However, they still mostly utilize the classical Byte-Pair Encoding (BPE) algorithm, a simple and reliable method. BPE has been shown to cause such issues as under-trained tokens and sub-optimal compression that may affect the downstre…

2024

Multilinguality or Back-translation? A Case Study with Estonian

COLING 2024main

Machine translation quality is highly reliant on large amounts of training data, and, when a limited amount of parallel data is available, synthetic back-translated or multilingual data can be used in addition. In this work, we introduce SynEst, a synthetic corpus of translations from 11 languages i…

Cited by 0SourcePDFScholar