MathCritique: Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Training critique models to provide useful feedback for actor models is an effective approach in scalable oversight, especially for complex tasks like math reasoning. However, current research lacks suitable datasets for effectively training critique models and integrating them in a principled way a