How a Bilingual LM Becomes Bilingual: Tracing Internal Representations with Sparse Autoencoders
This study explores how bilingual language models develop complex internal representations.We employ sparse autoencoders to analyze internal representations of bilingual language models with a focus on the effects of training steps, layers, and model sizes.Our analysis shows that language models fir