← Search

Fan Shu

1 accepted papers

2026

DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science

ICLR 2026poster

The fast-growing demands in using Large Language Models (LLMs) to tackle complex multi-step data science tasks create a emergent need for accurate benchmarking. There are two major gaps in existing benchmarks: (i) the lack of standardized, process-aware evaluation that captures instruction adherence…

Cited by 0SourcecodeScholar