← Search

Hongru Yang

3 accepted papers

2025

Transformers Provably Learn Two-Mixture of Linear Classification via Gradient Flow

ICLR 2025poster

Understanding how transformers learn and utilize hidden connections between tokens is crucial to understand the behavior of large language models. To understand this mechanism, we consider the task of two-mixture of linear classification which possesses a hidden correspondence structure among tokens…

Cited by 0SourcePDFScholar
2024

Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis

NeurIPS 2024poster

Understanding the training dynamics of transformers is important to explain the impressive capabilities behind large language models. In this work, we study the dynamics of training a shallow transformer on a task of recognizing co-occurrence of two designated words. In the literature of studying t…

Cited by 1SourcePDFScholar