How Transformers Learn Causal Structures In-Context: Explainable Mechanism Meets Theoretical Guarantee
Transformers have demonstrated remarkable in-context learning abilities, adapting to new tasks from just a few examples without parameter updates. However, theoretical understanding of this phenomenon typically assumes fixed dependency structures, while real-world sequences exhibit flexible, context…