← Search

Aleksandr Shevchenko

4 accepted papers

2026

Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models

ICLR 2026poster

Recent empirical studies have explored the idea of continuing to train a model at test-time for a given task, known as test-time training (TTT), and have found it to yield significant performance improvements. However, there is limited understanding of why and when TTT is effective. Earlier explanat…

Cited by 0SourceScholar
2025

Attention with Trained Embeddings Provably Selects Important Tokens

NeurIPS 2025poster

Token embeddings play a crucial role in language modeling but, despite this practical relevance, their theoretical understanding is limited. Our paper addresses the gap by characterizing the structure of embeddings obtained via gradient descent. Specifically, we consider a one-layer softmax attentio…

Cited by 0SourceScholar
2024

Compression of Structured Data with Autoencoders: Provable Benefit of Nonlinearities and Depth

ICML 2024poster

Autoencoders are a prominent model in many empirical branches of machine learning and lossy data compression. However, basic theoretical questions remain unanswered even in a shallow two-layer setting. In particular, to what degree does a shallow autoencoder capture the structure of the underlying d…

Cited by 4SourcePDFScholar
2023

Fundamental Limits of Two-layer Autoencoders, and Achieving Them with Gradient Methods

ICML 2023oral

Autoencoders are a popular model in many branches of machine learning and lossy data compression. However, their fundamental limits, the performance of gradient methods and the features learnt during optimization remain poorly understood, even in the two-layer setting. In fact, earlier work has cons…

Cited by 8SourcePDFScholar