← Search

Yash Sarrof

4 accepted papers

2026

On the Ability of Transformers to Verify Plans

ICML 2026poster

Transformers have shown inconsistent success in AI planning tasks, and theoretical understanding of when generalization should be expected has been limited. We take important steps towards addressing this gap by analyzing the ability of decoder-only models to verify whether a given plan correctly so…

Cited by 0SourceScholar
2025

A Formal Framework for Understanding Length Generalization in Transformers

ICLR 2025poster

A major challenge for transformers is generalizing to sequences longer than those observed during training. While previous works have empirically shown that transformers can either succeed or fail at length generalization depending on the task, theoretical understanding of this phenomenon remains li…

2025

Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities

NeurIPS 2025poster

Transformers have theoretical limitations in modeling certain sequence-to-sequence tasks, yet it remains largely unclear if these limitations play a role in large-scale pretrained LLMs, or whether LLMs might effectively overcome these constraints in practice due to the scale of both the models thems…

Cited by 0SourceScholar
2024

The Expressive Capacity of State Space Models: A Formal Language Perspective

NeurIPS 2024poster

Recently, recurrent models based on linear state space models (SSMs) have shown promising performance in language modeling (LM), competititve with transformers. However, there is little understanding of the in-principle abilities of such models, which could provide useful guidance to the search for…

Cited by 3SourcePDFScholar