← Search

Sam Bowyer

2 accepted papers

2025

Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints

ICML 2025spotlight

Rigorous statistical evaluations of large language models (LLMs), including valid error bars and significance testing, are essential for meaningful and reliable performance assessment. Currently, when such statistical measures are reported, they typically rely on the Central Limit Theorem (CLT). In…

Cited by 17SourcePDFScholar
2024

Using Autodiff to Estimate Posterior Moments, Marginals and Samples

UAI 2024poster

Importance sampling is a popular technique in Bayesian inference: by reweighting samples drawn from a proposal distribution we are able to obtain samples and moment estimates from a Bayesian posterior over latent variables. Recent work, however, indicates that importance sampling scales poorly — in…