← Search

Shahar Katz

6 accepted papers

2024

Backward Lens: Projecting Language Model Gradients into the Vocabulary Space

EMNLP 2024main

Understanding how Transformer-based Language Models (LMs) learn and recall information is a key goal of the deep learning community. Recent interpretability methods project weights and hidden states obtained from the forward pass to the models’ vocabularies, helping to uncover how information flows…