← Search

Michael R DeWeese

4 accepted papers

2026

Bits That Count: Quantifying and Predicting Capabilities of Language Models

ICML 2026poster

What and how do language models learn during training? When does learning elicit \textit{existing} knowledge, and when does it primarily teach \textit{new} capabilities? We find that the amount of generalizable information language models learn during training predicts the origins of their emergent …

Cited by 0SourceScholar
2025

Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks

NeurIPS 2025poster

What features neural networks learn, and how, remains an open question. In this paper, we introduce Alternating Gradient Flows (AGF), an algorithmic framework that describes the dynamics of feature learning in two-layer networks trained from small initialization. Prior works have shown that gradient…

Cited by 0SourceScholar
2025

Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like Models

NeurIPS 2025poster

Self-supervised word embedding algorithms such as word2vec provide a minimal setting for studying representation learning in language modeling. We examine the quartic Taylor approximation of the word2vec loss around the origin, and we show that both the resulting training dynamics and the final perf…

Cited by 0SourceScholar
2025

Quantifying Elicitation of Latent Capabilities in Language Models

NeurIPS 2025poster

Large language models often possess latent capabilities that lie dormant unless explicitly elicited, or surfaced, through fine-tuning or prompt engineering. Predicting, assessing, and understanding these latent capabilities pose significant challenges in the development of effective, safe AI systems…

Cited by 0SourceScholar