Beyond Step Pruning: Information Theory Based Step-level Optimization for Self-Refining Large Language Models
Large language models (LLMs) have shown impressive capabilities in natural language tasks, yet they continue to struggle with multi-step mathematical reasoning, where correctness depends on a precise chain of intermediate steps. Preference optimization methods such as Direct Preference Optimization