← Search

Tankred Saanum

4 accepted papers

2025

Sparse Autoencoders Reveal Temporal Difference Learning in Large Language Models

ICLR 2025poster

In-context learning, the ability to adapt based on a few examples in the input prompt, is a ubiquitous feature of large language models (LLMs). However, as LLMs' in-context learning abilities continue to improve, understanding this phenomenon mechanistically becomes increasingly important. In partic…

Cited by 7SourcePDFScholar
2024

Evaluating alignment between humans and neural network representations in image-based learning tasks

NeurIPS 2024poster

Humans represent scenes and objects in rich feature spaces, carrying information that allows us to generalise about category memberships and abstract functions with few examples. What determines whether a neural network model generalises like a human? We tested how well the representations of $86$ p…

2023

Reinforcement Learning with Simple Sequence Priors

NeurIPS 2023poster

In reinforcement learning (RL), simplicity is typically quantified on an action-by-action basis -- but this timescale ignores temporal regularities, like repetitions, often present in sequential strategies. We therefore propose an RL algorithm that learns to solve tasks with sequences of actions tha…

Cited by 26SourcePDFScholar