2026
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
ICML 2026poster
Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information during an auditing process, we introduce *overthinking*: the process of using reasoning task vectors to amplify the chai…