← Search

Geonhee Kim

2 accepted papers

2026

Mitigating Content Effects on Reasoning in Language Models Through Fine-Grained Activation Steering

AAAI 2026technical

Large language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, where plausible arguments are incorrectly deemed logically valid or vice versa. This paper investigates how content biases

Cited by 0SourcePDFScholar
2025

Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference

ACL 2025finding

Recent studies on reasoning in language models (LMs) have sparked a debate on whether they can learn systematic inferential principles or merely exploit superficial patterns in the training data. To understand and uncover the mechanisms adopted for formal reasoning in LMs, this paper presents a mech…

Cited by 0SourcePDFScholar