2026
Subword Tokenization for Low- and Medium-Resource Languages: A Systematic Evaluation
IJCAI 2026
Subword tokenization is a standard technique for pre-trained language models, mapping text into sequences of tokens from a fixed-size vocabulary. Despite its widespread use, the impact of tokenization algorithms and vocabulary sizes on downstream performance remains underexplored, particularly for l