← Search

Kaya Stechly

4 accepted papers

2026

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!

ICML 2026poster

Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the performance of language models on reasoning tasks. These intermediate tokens have been called \say{reasoning traces} or even \say{thoughts} -- implicitly anthropomorphi…

Cited by 39SourceScholar
2025

On the self-verification limitations of large language models on reasoning and planning tasks

ICLR 2025poster

There has been considerable divergence of opinion on the reasoning abilities of Large Language Models (LLMs). While the initial optimism that reasoning might emerge automatically with scale has been tempered thanks to a slew of counterexamples--ranging from multiplication to simple planning--there p…

Cited by 57SourcePDFScholar
2024

Chain of Thoughtlessness? An Analysis of CoT in Planning

NeurIPS 2024poster

Large language model (LLM) performance on reasoning problems typically does not generalize out of distribution. Previous work has claimed that this can be mitigated with chain of thought prompting--a method of demonstrating solution procedures--with the intuition that it is possible to in-context te…

Cited by 47SourcePDFScholar
2024

Position: LLMs Can’t Plan, But Can Help Planning in LLM-Modulo Frameworks

ICML 2024spotlight

We argue that auto-regressive LLMs cannot, by themselves, do planning or self-verification (which is after all a form of reasoning), and shed some light on the reasons for misunderstandings in the literature. We will also argue that LLMs should be viewed as universal approximate knowledge sources th…

Cited by 192SourcePDFScholar