2025
Pre-trained Models Perform the Best When Token Distributions Follow Zipf’s Law
EMNLP 2025
Tokenization is a fundamental step in natural language processing (NLP) and other sequence modeling domains, where the choice of vocabulary size significantly impacts model performance. Despite its importance, selecting an optimal vocabulary size remains underexplored, typically relying on heuristic