← Search

Praneet Suresh

3 accepted papers

2026

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

ICML 2026poster

Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face unexpected or adversarial data that diverges from training data distributions. Without explicit mechanisms for handling s…

Cited by 0SourceScholar
2026

Quantifying LLM Attention-Head Stability: Implications for Circuit Universality

ICML 2026poster

In mechanistic interpretability, recent work scrutinizes transformer “circuits”—sparse, mono or multi layer sub computations, that may reflect human understandable functions. Yet, these network circuits are rarely acid-tested for their stability across different instances of the same deep learning a…

Cited by 0SourceScholar
2025

From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers

NeurIPS 2025poster

As generative AI systems become competent and democratized in science, business, and government, deeper insight into their failure modes now poses an acute need. The occasional volatility in their behavior, such as the propensity of transformer models to hallucinate, impedes trust and adoption of em…

Cited by 0SourceScholar