← Search

Ziqian Zhong

4 accepted papers

2026

ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases

ICLR 2026poster

The tendency to find and exploit "shortcuts" to complete tasks poses significant risks for reliable assessment and deployment of large language models (LLMs). For example, an LLM agent with access to unit tests may delete failing tests rather than fix the underlying bug. Such behavior undermines bot…

Cited by 0SourcecodeScholar
2023

The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks

NeurIPS 2023oral

Do neural networks, trained on well-understood algorithmic tasks, reliably rediscover known algorithms? Several recent studies, on tasks ranging from group operations to in-context linear regression, have suggested that the answer is yes. Using modular addition as a prototypical problem, we show tha…