← Search

Qinan Yu

7 accepted papers

2026

Outcome-Based Rewards Do Not Guarantee Faithful and Verifiable Reasoning

ICML 2026poster

Reinforcement Learning from Verifiable Rewards (RLVR) on chain-of-thought reasoning has become a standard part of language model post-training recipes. A common assumption is that the reasoning chains trained through RLVR represent how a model gets to its answer. In this paper, we develop two metric…

Cited by 0SourceScholar
2025

Improved Representation Steering for Language Models

NeurIPS 2025spotlight

Steering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behavior. Recent work has shown that adjusting weights or representations is often less effective than st…

Cited by 0SourcecodeScholar
2025

The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling

ICLR 2025poster

We employ new tools from mechanistic interpretability to ask whether the internal structure of large language models (LLMs) shows correspondence to the linguistic structures which underlie the languages on which they are trained. In particular, we ask (1) when two languages employ the same morphosyn…

Cited by 2SourcePDFScholar
2024

LLM Circuit Analyses Are Consistent Across Training and Scale

NeurIPS 2024poster

Most currently deployed LLMs undergo continuous training or additional finetuning. By contrast, most research into LLMs' internal mechanisms focuses on models at one snapshot in time (the end of pre-training), raising the question of whether their results generalize to real-world settings. Existing…

Cited by 8SourcePDFScholar
2023

Are Language Models Worse than Humans at Following Prompts? It's Complicated

EMNLP 2023short findings

Prompts have been the center of progress in advancing language models' zero-shot and few-shot performance. However, recent work finds that models can perform surprisingly well when given intentionally irrelevant or misleading prompts. Such results may be interpreted as evidence that model behavior i…

Cited by 0SourcecodeScholar