← Search

Giordano Rogers

2 accepted papers

2026

Do Natural Language Interpretability Methods Convey Privileged Information?

ICML 2026poster

Recent interpretability methods have proposed to translate LLM internal representations into natural language descriptions using a second verbalizer LLM. This is intended to illuminate how the target model represents and operates on inputs. But do such activation verbalization approaches actually pr…

Cited by 0SourceScholar