← Search

Mikail Khona

8 accepted papers

2025

Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing

ICML 2025poster

Knowledge Editing (KE) algorithms alter models' weights to perform targeted updates to incorrect, outdated, or otherwise unwanted factual associations. However, recent work has shown that applying KE can adversely affect models' broader factual recall accuracy and diminish their reasoning abilities.…

Cited by 0SourcePDFScholar
2025

Uncovering Latent Memories in Large Language Models

ICLR 2025poster

Frontier AI systems are making transformative impacts across society, but such benefits are not without costs: models trained on web-scale datasets containing personal and private data raise profound concerns about data privacy and security. Language models are trained on extensive corpora including…

Cited by 0SourcePDFScholar
2024

Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks

ICML 2024poster

Transformers trained on huge text corpora exhibit a remarkable set of capabilities, e.g., performing simple logical operations. Given the inherent compositional nature of language, one can expect the model to learn to compose these capabilities, potentially yielding a combinatorial explosion of what…

2024

Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model

ICML 2024poster

Stepwise inference protocols, such as scratchpads and chain-of-thought, help language models solve complex problems by decomposing them into a sequence of simpler subproblems. To unravel the underlying mechanisms of stepwise inference we propose to study autoregressive Transformer models on a synthe…

Cited by 4SourcePDFScholar
2023

Self-Supervised Learning of Representations for Space Generates Multi-Modular Grid Cells

NeurIPS 2023poster

To solve the spatial problems of mapping, localization and navigation, the mammalian lineage has developed striking spatial representations. One important spatial representation is the Nobel-prize winning grid cells: neurons that represent self-location, a local and aperiodic quantity, with seemingl…

Cited by 22SourcePDFScholar
2022

No Free Lunch from Deep Learning in Neuroscience: A Case Study through Models of the Entorhinal-Hippocampal Circuit

NeurIPS 2022accept

Research in Neuroscience, as in many scientific disciplines, is undergoing a renaissance based on deep learning. Unique to Neuroscience, deep learning models can be used not only as a tool but interpreted as models of the brain. The central claims of recent deep learning-based models of brain circui…

Cited by 73SourcePDFScholar
2021

Efficient online inference for nonparametric mixture models

UAI 2021poster

Natural data are often well-described as belonging to latent clusters. When the number of clusters is unknown, Bayesian nonparametric (BNP) models can provide a flexible and powerful technique to model the data. However, algorithms for inference in nonparametric mixture models fail to meet two criti…

2020

Reverse-engineering recurrent neural network solutions to a hierarchical inference task for mice

NeurIPS 2020poster

We study how recurrent neural networks (RNNs) solve a hierarchical inference task involving two latent variables and disparate timescales separated by 1-2 orders of magnitude. The task is of interest to the International Brain Laboratory, a global collaboration of experimental and theoretical neuros…

Cited by 41SourcePDFScholar