← Search

Benjamin L. Edelman

8 accepted papers

2024

Distinguishing the Knowable from the Unknowable with Language Models

ICML 2024poster

We study the feasibility of identifying *epistemic* uncertainty (reflecting a lack of knowledge), as opposed to *aleatoric* uncertainty (reflecting entropy in the underlying distribution), in the outputs of large language models (LLMs) over free-form text. In the absence of ground-truth probabilitie…

2024

Feature emergence via margin maximization: case studies in algebraic tasks

ICLR 2024spotlight

Understanding the internal representations learned by neural networks is a cornerstone challenge in the science of machine learning. While there have been significant recent strides in some cases towards understanding *how* neural networks implement specific target functions, this paper explores a c…

Cited by 15SourcePDFScholar
2024

The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains

NeurIPS 2024poster

Large language models have the ability to generate text that mimics patterns in their inputs. We introduce a simple Markov Chain sequence modeling task in order to study how this in-context learning capability emerges. In our setting, each example is sampled from a Markov chain drawn from a prior di…

Cited by 41SourcePDFScholar
2024

Transcendence: Generative Models Can Outperform The Experts That Train Them

NeurIPS 2024poster

Generative models are trained with the simple objective of imitating the conditional probability distribution induced by the data they are trained on. Therefore, when trained on data generated by humans, we may not expect the artificial model to outperform the humans on their original objectives. In…

Cited by 11SourcePDFScholar
2024

Watermarks in the Sand: Impossibility of Strong Watermarking for Language Models

ICML 2024poster

Watermarking generative models consists of planting a statistical signal (watermark) in a model's output so that it can be later verified that the output was generated by the given model. A strong watermarking scheme satisfies the property that a computationally bounded attacker cannot erase the wat…

Cited by 6SourcePDFScholar
2023

Pareto Frontiers in Deep Feature Learning: Data, Compute, Width, and Luck

NeurIPS 2023spotlight

In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs. This work investigates how these complexities necessarily arise for feature learning in the presence of computational-statistical gaps. We begin by considering off…

Cited by 4SourcePDFScholar
2022

Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit

NeurIPS 2022accept

There is mounting evidence of emergent phenomena in the capabilities of deep learning methods as we scale up datasets, model sizes, and training times. While there are some accounts of how these resources modulate statistical capacity, far less is known about their effect on the computational proble…

Cited by 156SourcePDFScholar
2022

Inductive Biases and Variable Creation in Self-Attention Mechanisms

ICML 2022spotlight

Self-attention, an architectural motif designed to model long-range interactions in sequential data, has driven numerous recent breakthroughs in natural language processing and beyond. This work provides a theoretical analysis of the inductive biases of self-attention modules. Our focus is to rigoro…

Cited by 159SourcePDFScholar