← Search

Yanjin He

1 accepted papers

2025

Pre-trained Models Perform the Best When Token Distributions Follow Zipf’s Law

EMNLP 2025

Tokenization is a fundamental step in natural language processing (NLP) and other sequence modeling domains, where the choice of vocabulary size significantly impacts model performance. Despite its importance, selecting an optimal vocabulary size remains underexplored, typically relying on heuristic