When Reasoning Collapses: A Depth-Aware Probe into LLM Reasoning (Student Abstract)
Large language models (LLMs) often perform better when prompted to explain their reasoning, but it remains unclear how well such gains persist as reasoning depth increases. In this work, we propose a depth-aware evaluation framework alongside the performance results on two structured datasets: CLUTR