← Search

Tom Kempton

3 accepted papers

2026

DMAP: A Distribution Map for Text

ICLR 2026poster

Large Language Models (LLMs) are a powerful tool for statistical text analysis, with derived sequences of next-token probability distributions offering a wealth of information. Extracting this signal typically relies on metrics such as perplexity, which do not adequately account for context; how one…

Cited by 0SourceScholar
2025

Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Models

EMNLP 2025

Advances in hardware and language model architecture have spurred a revolution in natural language generation. However, autoregressive models compute probability distributions over next-token choices, and sampling from these distributions, known as decoding, has received significantly less attention

2025

TempTest: Local Normalization Distortion and the Detection of Machine-generated Text

AISTATS 2025poster

Existing methods for the zero-shot detection of machine-generated text are dominated by three statistical quantities: log-likelihood, log rank, and entropy. As language models mimic the distribution of human text ever closer, this will limit our ability to build effective detection algorithms. To co…

Cited by 0SourcecodeScholar