AAAI 2026technical0 citations
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
Andrey V. Galichin, Alexey Dontsov, Polina Druzhinina, Anton Razzhigaev, Oleg Rogov, Elena Tutubalina, Ivan Oseledets
Abstract
Recent LLMs like DeepSeek-R1 have demonstrated state-of-the-art performance by integrating deep thinking and complex reasoning during generation. However, the internal mechanisms behind these reasoning processes remain unexplored. We observe reasoning LLMs consistently use vocabulary associated with human reasoning processes. We hypothesize these words correspond to specific reasoning moments within the models
BibTeX
@inproceedings{aaai2026_ihavecoveredallt,
title = {I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders},
author = {Andrey V. Galichin and Alexey Dontsov and Polina Druzhinina and Anton Razzhigaev and Oleg Rogov and Elena Tutubalina and Ivan Oseledets},
booktitle = {AAAI 2026},
year = {2026}
}