2024
Direct Evaluation of Chain-of-Thought in Multi-hop Reasoning with Knowledge Graphs
ACL 2024findings
Large language models (LLMs) have demonstrated strong reasoning abilities when prompted to generate chain-of-thought (CoT) explanations alongside answers. However, previous research on evaluating LLMs has solely focused on answer accuracy, neglecting the correctness of the generated CoT. In this pap…