← Search

Jasper Lee

2 accepted papers

2025

CLEVER: A Curated Benchmark for Formally Verified Code Generation

NeurIPS 2025poster

We introduce ${\rm C{\small LEVER}}$, a high-quality, manually curated benchmark of 161 problems for end-to-end verified code generation in Lean. Each problem consists of (1) the task of generating a specification that matches a held-out ground-truth specification, and (2) the task of generating a L…

Cited by 0SourcecodeScholar
2024

PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition

NeurIPS 2024poster

We present PutnamBench, a new multi-language benchmark for evaluating the ability of neural theorem-provers to solve competition mathematics problems. PutnamBench consists of 1692 hand-constructed formalizations of 640 theorems sourced from the William Lowell Putnam Mathematical Competition, the pre…