Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
Despite rapid advances in code generation, current Large Language Models (LLMs) still lack an essential capability for reliable and verifiable code generation: compositional reasoning across multi-function programs. To explore this potential and important gap, we introduce DafnyCOMP, a benchmark des…