← Search

Sangwhan Moon

2 accepted papers

2026

Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models

ICML 2026poster

Byte-level tokenization enables language models to handle any Unicode input, but models can generate invalid UTF-8 sequences when encountering rare or unseen characters. We investigate the relationship between training scale and UTF-8 generation reliability with a 355M parameter model trained on 80B…

Cited by 0SourceScholar
2024

Two Counterexamples to Tokenization and the Noiseless Channel

COLING 2024main

In Tokenization and the Noiseless Channel (Zouhar et al., 2023), Rényi efficiency is suggested as an intrinsic mechanism for evaluating a tokenizer: for NLP tasks, the tokenizer which leads to the highest Rényi efficiency of the unigram distribution should be chosen. The Rényi efficiency is thus tre…