NeurIPS 2025poster0 citations

A unified framework for establishing the universal approximation of transformer-type architectures

Jingpu Cheng, Ting Lin, Zuowei Shen, Qianxiao Li

Abstract

We investigate the universal approximation property (UAP) of transformer-type architectures, providing a unified theoretical framework that extends prior results on residual networks to models incorporating attention mechanisms. Our work identifies token distinguishability as a fundamental requirement for UAP and introduces a general sufficient condition that applies to a broad class of architectures. Leveraging an analyticity assumption on the attention layer, we can significantly simplify the verification of this condition, providing a non-constructive approach in establishing UAP for such architectures. We demonstrate the applicability of our framework by proving UAP for transformers with various attention mechanisms, including kernel-based and sparse ones. The corollaries of our results either generalize prior works or establish UAP for architectures not previously covered. Furthermore, our framework offers a principled foundation for designing novel transformer architectures with inherent UAP guarantees, including those with specific functional symmetries. We propose examples to illustrate these insights.

transformersapproximation theoryequivariance
BibTeX
@inproceedings{
cheng2025a,
title={A unified framework for establishing the universal approximation of transformer-type architectures},
author={Jingpu Cheng and Ting Lin and Zuowei Shen and Qianxiao Li},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=2xjcosH3yQ}
}