← Search

Kshitij Dubey

1 accepted papers

2025

RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation

ICML 2025poster

Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true “reasoning” or from statistical recall of the training set. Inspired by the ladder of causation (Pearl, 2009) and its three levels (assoc…

Cited by 0SourcePDFScholar