← Search

Samet Demir

3 accepted papers

2026

Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions

ICML 2026poster

Pretrained Transformers can perform in-context learning (ICL) from a few demonstrations, but this ability can fail sharply when the test distribution differs from pretraining—a common deployment setting. We study attention temperature as a simple inference-time control for improving ICL robustness u…

Cited by 0SourceScholar
2025

Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure

ICLR 2025poster

In this work, we study the training and generalization performance of two-layer neural networks (NNs) after one gradient descent step under structured data modeled by Gaussian mixtures. While previous research has extensively analyzed this model under isotropic data assumption, such simplifications…

2025

How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs

NeurIPS 2025poster

Pretrained Transformers demonstrate remarkable in-context learning (ICL) capabilities, enabling them to adapt to new tasks from demonstrations without parameter updates. However, theoretical studies often rely on simplified architectures (e.g., omitting MLPs), data models (e.g., linear regression wi…

Cited by 0SourceScholar