2025
Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
NAACL 2025long
Self-improvement methods enable large language models (LLMs) to generate solutions themselves and iteratively train on filtered, high-quality rationales. This process proves effective and reduces the reliance on human supervision in LLMs’ reasoning, but the performance soon plateaus. We delve into t…