2026
HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
ICLR 2026poster
While Large Language Models (LLMs) have demonstrated significant advancements in reasoning and agent-based problem-solving, current evaluation methodologies fail to adequately assess their capabilities: existing benchmarks either rely on closed-ended questions prone to saturation and memorization, o…