← Search

Nuri Mert Vural

2 accepted papers

2026

Learning to Recall with Transformers Beyond Orthogonal Embeddings

ICLR 2026poster

Modern large language models (LLMs) excel at tasks that require storing and retrieving knowledge, such as factual recall and question answering. Transformers are central to this capability, thanks to their ability to encode information during training and retrieve it at inference. Existing theoretic…

Cited by 0SourceScholar
2025

Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws

NeurIPS 2025poster

We study the optimization and sample complexity of gradient-based training of a two-layer neural network with quadratic activation function in the high-dimensional regime, where the data is generated as $y \propto \sum_{j=1}^{r}\lambda_j \sigma\left(\langle \boldsymbol{\theta_j}, \boldsymbol{x}\rang…

Cited by 0SourceScholar