← Search

Aditya V. Nori

5 accepted papers

2025

Compositional Causal Reasoning Evaluation in Language Models

ICML 2025poster

Causal reasoning and compositional reasoning are two core aspirations in AI. Measuring the extent of these behaviors requires principled evaluation methods. We explore a unified perspective that considers both behaviors simultaneously, termed *compositional causal reasoning* (CCR): the ability to in…

Cited by 1SourcePDFScholar
2025

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation

ICML 2025poster

Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true “reasoning” or from statistical recall of the training set. Inspired by the ladder of causation (Pearl, 2009) and its three levels (assoc…

Cited by 0SourcePDFScholar
2025

Reasoning Elicitation in Language Models via Counterfactual Feedback

ICLR 2025oral

Despite the increasing effectiveness of language models, their reasoning capabilities remain underdeveloped. In particular, causal reasoning through counterfactual question answering is lacking. This work aims to bridge this gap. We first derive novel metrics that balance accuracy in factual and cou…

Cited by 0SourcePDFScholar
2024

Does Reasoning Emerge? Examining the Probabilities of Causation in Large Language Models

NeurIPS 2024poster

Recent advances in AI have been significantly driven by the capabilities of large language models (LLMs) to solve complex problems in ways that resemble human thinking. However, there is an ongoing debate about the extent to which LLMs are capable of actual reasoning. Central to this debate are two…

Cited by 3SourcePDFScholar
2023

Exploring the Boundaries of GPT-4 in Radiology

EMNLP 2023long main

The recent success of general-domain large language models (LLMs) has significantly changed the natural language processing paradigm towards a unified foundation model across domains and applications. In this paper, we focus on assessing the performance of GPT-4, the most capable LLM so far, on the…

Cited by 0SourceScholar