← Search

Jingpu Cheng

5 accepted papers

2025

A unified framework for establishing the universal approximation of transformer-type architectures

NeurIPS 2025poster

We investigate the universal approximation property (UAP) of transformer-type architectures, providing a unified theoretical framework that extends prior results on residual networks to models incorporating attention mechanisms. Our work identifies token distinguishability as a fundamental requireme…

Cited by 0SourceScholar
2025

From Weight-Based to State-Based Fine-Tuning: Further Memory Reduction on LoRA with Parallel Control

ICML 2025oral

The LoRA method has achieved notable success in reducing GPU memory usage by applying low-rank updates to weight matrices. Yet, one simple question remains: can we push this reduction even further? Furthermore, is it possible to achieve this while improving performance and reducing computation time?…

Cited by 0SourcePDFScholar
2024

An Optimal Control View of LoRA and Binary Controller Design for Vision Transformers

ECCV 2024poster

"While recent advancements in model fine-tuning predominantly emphasize the utilization of low-rank adaptation (LoRA), we propose an alternative approach centered on reducing the precision of adaptation matrices. In particular, we depart from the common viewpoint that considers adaptation matrices s…

Cited by 0SourcePDFScholar