ICLR 2026poster0 citations

On learning linear dynamical systems in context with attention layers

Maria-Luiza Vladarean, Xuhui Zhang, Suvrit Sra

Abstract

This paper studies the expressive power of linear attention layers for in-context learning (ICL) of linear dynamical systems (LDS). We consider training on sequences of inexact observations produced by noise-corrupted LDSs, with all perturbations being Gaussian; importantly, we study the non-i.i.d. setting as it is closer to real-world scenarios. We provide the optimal weight construction for a single linear-attention layer and show its equivalence to one step of Gradient Descent relative to an autoregression objective of window size one. Guided by experiments, we uncover a relation to the Preconditioned Conjugate Gradient method for larger window sizes. We back our findings with numerical evidence. These results add to the existing understanding of transformers' expressivity as in-context learners, and offer plausible hypotheses for experimental observations whereby they compete with Kalman filters --- the optimal model-dependent learners for this setting.

in-context learninglinear attentionlinear dynamical systemskalman filtertime series
BibTeX
@inproceedings{
vladarean2026on,
title={On learning linear dynamical systems in context with attention layers},
author={Maria-Luiza Vladarean and Xuhui Zhang and Suvrit Sra},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=os7OLubIMI}
}