2023
DialCoT Meets PPO: Decomposing and Exploring Reasoning Paths in Smaller Language Models
EMNLP 2023long main
Chain-of-Thought (CoT) prompting has successfully enhanced the reasoning capabilities of Large Language Models~(LLMs) with at least 100 billion parameters. However, it is ineffective, or even detrimental, to the performance on reasoning tasks in Smaller Language Models (SLMs) with less than 10 billi…