CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation in Real-World API Interactions
Assessing the capabilities of large language models (LLMs) as agents in decision making and operational tasks is crucial for the development of LLM-as-agent service. We propose CToolEval, a benchmark designed to evaluate LLMs in the context of Chinese societal applications, featuring 398 APIs across…