← Search

Elan Rosenfeld

14 accepted papers

2026

Deep sequence models tend to memorize geometrically; it is unclear why.

ICML 2026poster

Deep sequence models are said to store atomic facts predominantly in the form of associative memory: a brute-force lookup of co-occurring entities. We identify a dramatically different form of storage of atomic facts that we term as geometric memory. Here, the model has synthesized embeddings encodi…

Cited by 0SourceScholar
2024

Identifying Representations for Intervention Extrapolation

ICLR 2024poster

The premise of identifiable and causal representation learning is to improve the current representation learning paradigm in terms of generalizability or robustness. Despite recent progress in questions of identifiability, more theoretical results demonstrating concrete advantages of these methods f…

Cited by 18SourcePDFScholar
2024

Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization

ICLR 2024poster

We identify a new phenomenon in neural network optimization which arises from the interaction of depth and a particular heavy-tailed structure in natural data. Our result offers intuitive explanations for several previously reported observations about network training dynamics, including a conceptua…

Cited by 11SourcePDFScholar
2023

(Almost) Provable Error Bounds Under Distribution Shift via Disagreement Discrepancy

NeurIPS 2023poster

We derive a new, (almost) guaranteed upper bound on the error of deep neural networks under distribution shift using unlabeled test data. Prior methods are either vacuous in practice or accurate on average but heavily underestimate error for a sizeable fraction of shifts. In particular, the latter o…

2023

Learning Linear Causal Representations from Interventions under General Nonlinear Mixing

NeurIPS 2023oral

We study the problem of learning causal representations from unknown, latent interventions in a general setting, where the latent distribution is Gaussian but the mixing function is completely general. We prove strong identifiability results given unknown single-node interventions, i.e., without hav…

Cited by 71SourcePDFScholar
2022

An Online Learning Approach to Interpolation and Extrapolation in Domain Generalization

AISTATS 2022poster

A popular assumption for out-of-distribution generalization is that the training data comprises sub-datasets, each drawn from a distinct distribution; the goal is then to "interpolate" these distributions and "extrapolate" beyond them—this objective is broadly known as domain generalization. A commo…

Cited by 36SourcePDFScholar
2022

Analyzing and Improving the Optimization Landscape of Noise-Contrastive Estimation

ICLR 2022spotlight

Noise-contrastive estimation (NCE) is a statistically consistent method for learning unnormalized probabilistic models. It has been empirically observed that the choice of the noise distribution is crucial for NCE’s performance. However, such observation has never been made formal or quantitative. I…

Cited by 22SourcePDFScholar
2022

Iterative Feature Matching: Toward Provable Domain Generalization with Logarithmic Environments

NeurIPS 2022accept

Domain generalization aims at performing well on unseen test environments with data from a limited number of training environments. Despite a proliferation of proposed algorithms for this task, assessing their performance both theoretically and empirically is still very challenging. Distributional m…

Cited by 43SourcePDFScholar
2020

Certified Robustness to Label-Flipping Attacks via Randomized Smoothing

ICML 2020poster

Machine learning algorithms are known to be susceptible to data poisoning attacks, where an adversary manipulates the training data to degrade performance of the resulting classifier. In this work, we present a unifying view of randomized smoothing over arbitrary functions, and we leverage this nove…

Cited by 214SourcePDFScholar