AAAI 2026technical0 citations

I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders

Andrey V. Galichin, Alexey Dontsov, Polina Druzhinina, Anton Razzhigaev, Oleg Rogov, Elena Tutubalina, Ivan Oseledets

Abstract

Recent LLMs like DeepSeek-R1 have demonstrated state-of-the-art performance by integrating deep thinking and complex reasoning during generation. However, the internal mechanisms behind these reasoning processes remain unexplored. We observe reasoning LLMs consistently use vocabulary associated with human reasoning processes. We hypothesize these words correspond to specific reasoning moments within the models

BibTeX
@inproceedings{aaai2026_ihavecoveredallt,
  title = {I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders},
  author = {Andrey V. Galichin and Alexey Dontsov and Polina Druzhinina and Anton Razzhigaev and Oleg Rogov and Elena Tutubalina and Ivan Oseledets},
  booktitle = {AAAI 2026},
  year = {2026}
}