2026
The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
ICLR 2026oral
Existing interpretability methods for Large Language Models (LLMs) often fall short by focusing on linear directions or isolated features, overlooking the high-dimensional, nonlinear, and relational geometry within model representations. This study focuses on how adversarial inputs systematically af…