← Search

Yuhan Zheng

1 accepted papers

2025

AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science

EMNLP 2025

Large language models (LLMs) are increasingly used to automate data analysis through executable code generation. Yet, data science tasks often admit multiple statistically valid solutions—for example, different modeling strategies—making it critical to understand the reasoning behind analyses, not j