← Search

Miles Turpin

2 accepted papers

2025

Looking Inward: Language Models Can Learn About Themselves by Introspection

ICLR 2025poster

Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g. thoughts and feelings) that are not accessible to external observers. Do LLMs have this introspective capability of privileged acce…

2023

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

NeurIPS 2023poster

Large Language Models (LLMs) can achieve strong performance on many tasks by producing step-by-step reasoning before giving a final output, often referred to as chain-of-thought reasoning (CoT). It is tempting to interpret these CoT explanations as the LLM's process for solving a task. This level of…