← Search

Jiarui Jiang

2 accepted papers

2025

Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression

NeurIPS 2025poster

State-space models (SSMs), particularly Mamba, emerge as an efficient Transformer alternative with linear complexity for long-sequence modeling. Recent empirical works demonstrate Mamba's in-context learning (ICL) capabilities competitive with Transformers, a critical capacity for large foundation m…

Cited by 0SourceScholar
2024

Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization

NeurIPS 2024poster

Transformers have demonstrated great power in the recent development of large foundational models. In particular, the Vision Transformer (ViT) has brought revolutionary changes to the field of vision, achieving significant accomplishments on the experimental side. However, their theoretical capabili…

Cited by 5SourcePDFScholar