2026
Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
ICLR 2026poster
Polysemanticity is pervasive in language models and remains a major challenge for interpretation and model behavioral control. Leveraging sparse autoencoders (SAEs), we map the polysemantic topology of two small models (Pythia-70M and GPT-2-Small) to identify SAE feature pairs that are semantically…