2026
Scaling the Scaling Logic: Agentic Meta-Synthesis of Logic Reasoning
ICML 2026poster
Scaling verifiable training signals remains a key bottleneck for Reinforcement Learning from Verifiable Rewards (RLVR). Logical reasoning is a natural substrate: constraints are formal and answers are programmatically checkable. However, prior synthesis pipelines either depend on expert-written code…