2025
Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models
EMNLP 2025
Standard Byte-Pair Encoding (BPE) tokenization compresses text by pairing a learned token vocabulary with a detailed merge list. Recent work has shown that this merge list exposes a potential attack surface for extracting information about language model’s training data. In this paper, we explore th