← Search

Dhairya Dalal

2 accepted papers

2026

Mitigating Content Effects on Reasoning in Language Models Through Fine-Grained Activation Steering

AAAI 2026technical

Large language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, where plausible arguments are incorrectly deemed logically valid or vice versa. This paper investigates how content biases

Cited by 0SourcePDFScholar
2024

Inference to the Best Explanation in Large Language Models

ACL 2024long

While Large Language Models (LLMs) have found success in real-world applications, their underlying explanatory process is still poorly understood. This paper proposes IBE-Eval, a framework inspired by philosophical accounts on Inference to the Best Explanation (IBE) to advance the interpretation and…

Cited by 10SourcePDFScholar