WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models
Although Large Language Models (LLMs) excel in NLP tasks, they still need external tools to extend their ability. Current research on tool learning with LLMs often assumes mandatory tool use, which does not always align with real-world situations, where the necessity for tools is uncertain, and inco…