2026
FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
ICLR 2026poster
Large language models (LLMs) have recently demonstrated remarkable progress in formal theorem proving. Yet their ability to serve as practical assistants for mathematicians—filling in missing steps within complex proofs—remains underexplored. We identify this challenge as the task of subgoal complet…