← Search

Lucas Prieto

5 accepted papers

2026

Adversarial Vulnerability from Interference Between Features in Superposition

ICML 2026poster

Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-robust patterns in the input, and decision boundary structure, but none provides a representation-level mechanism that explains why specific perturbations succee…

Cited by 0SourceScholar
2026

Correlations in the Data Lead to Semantically Rich Feature Geometry Under Superposition

ICLR 2026poster

Recent advances in mechanistic interpretability have shown that many features represented by deep learning models can be captured by dictionary learning approaches such as sparse autoencoders. However, our understanding of the structures formed by these internal representations is still limited. Ini…

Cited by 0SourcecodeScholar
2026

Temporal superposition and feature geometry of RNNs under memory demands

ICLR 2026oral

Understanding how populations of neurons represent information is a central challenge across machine learning and neuroscience. Recent work in both fields has begun to characterize the representational geometry and functionality underlying complex distributed activity. For example, artificial neural…

Cited by 0SourceScholar
2025

Grokking at the Edge of Numerical Stability

ICLR 2025poster

Grokking, or sudden generalization that occurs after prolonged overfitting, is a surprising phenomenon that has challenged our understanding of deep learning. While a lot of progress has been made in understanding grokking, it is still not clear why generalization is delayed and why grokking often d…

2025

Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility

ICCV 2025poster

Robustness and resource-efficiency are two highly desirable properties for modern machine learning models. However, achieving them jointly remains a challenge. In this paper, we identify high learning rates as a facilitator for simultaneously achieving robustness to spurious correlations and network…