2025
Interpreting CLIP with Hierarchical Sparse Autoencoders
ICML 2025poster
Sparse autoencoders (SAEs) are useful for detecting and steering interpretable features in neural networks, with particular potential for understanding complex multimodal representations. Given their ability to uncover interpretable features, SAEs are particularly valuable for analyzing vision-langu…