← Search

George Ma

4 accepted papers

2026

Falsifying Sparse Autoencoder Reasoning Features in Language Models

ICML 2026poster

We study how reliably sparse autoencoders (SAEs) support claims about reasoning-related internal features in large language models. We first give a stylized analysis showing that sparsity-regularized decoding can preferentially retain stable low-dimensional correlates while suppressing high-dimensio…

Cited by 0SourceScholar
2024

A Canonicalization Perspective on Invariant and Equivariant Learning

NeurIPS 2024poster

In many applications, we desire neural networks to exhibit invariance or equivariance to certain groups due to symmetries inherent in the data. Recently, frame-averaging methods emerged to be a unified framework for attaining symmetries efficiently by averaging over input-dependent subsets of the gr…

2023

Laplacian Canonization: A Minimalist Approach to Sign and Basis Invariant Spectral Embedding

NeurIPS 2023poster

Spectral embedding is a powerful graph embedding technique that has received a lot of attention recently due to its effectiveness on Graph Transformers. However, from a theoretical perspective, the universal expressive power of spectral embedding comes at the price of losing two important invariance…