2026
Do Natural Language Interpretability Methods Convey Privileged Information?
ICML 2026poster
Recent interpretability methods have proposed to translate LLM internal representations into natural language descriptions using a second verbalizer LLM. This is intended to illuminate how the target model represents and operates on inputs. But do such activation verbalization approaches actually pr…