← Search

Aaryan Singhal

2 accepted papers

2025

LoLCATs: On Low-Rank Linearizing of Large Language Models

ICLR 2025poster

Recent works show we can linearize large language models (LLMs)—swapping the quadratic attentions of popular Transformer-based LLMs with subquadratic analogs, such as linear attention—avoiding the expensive pretraining costs. However, linearizing LLMs often significantly degrades model quality, stil…

2025

ThunderKittens: Simple, Fast, and $\textit{Adorable}$ Kernels

ICLR 2025spotlight

The challenge of mapping AI architectures to GPU hardware is creating a critical bottleneck in AI progress. Despite substantial efforts, hand-written custom kernels fail to meet their theoretical performance thresholds, even on well-established operations like linear attention. The diverse capabilit…

Cited by 0SourcePDFScholar