2025
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science
EMNLP 2025
Large language models (LLMs) are increasingly used to automate data analysis through executable code generation. Yet, data science tasks often admit multiple statistically valid solutions—for example, different modeling strategies—making it critical to understand the reasoning behind analyses, not j