Proxy-SPEX: Sample-Efficient Interpretability via Sparse Feature Interactions in LLMs
Large Language Models (LLMs) have achieved remarkable performance by capturing complex interactions between input features. To identify these interactions, most existing approaches require enumerating all possible combinations of features up to a given order, causing them to scale poorly with the nu…