← Search

Noa Nabeshima

3 accepted papers

2025

Learning Multi-Level Features with Matryoshka Sparse Autoencoders

ICML 2025poster

Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting neural networks by extracting the concepts represented in their activations. However, choosing the size of the SAE dictionary (i.e. number of learned concepts) creates a tension: as dictionary size increases to capture more…

2025

Parameterized Synthetic Text Generation with SimpleStories

NeurIPS 2025poster

We present SimpleStories, a large synthetic story dataset in simple language, consisting of 2 million samples each in English and Japanese. Through parameterizing prompts at multiple levels of abstraction, we achieve control over story characteristics at scale, inducing syntactic and semantic divers…

Cited by 0SourcecodeScholar
2022

Adversarial training for high-stakes reliability

NeurIPS 2022accept

In the future, powerful AI systems may be deployed in high-stakes settings, where a single failure could be catastrophic. One technique for improving AI safety in high-stakes settings is adversarial training, which uses an adversary to generate examples to train on in order to achieve better worst-c…

Cited by 64SourcePDFScholar