← Search

Mutsumi Nakamura

5 accepted papers

2025

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on tau-bench

EMNLP 2025

Recent advances in reasoning and planning capabilities of large language models (LLMs) have enabled their potential as autonomous agents capable of tool use in dynamic environments. However, in multi-turn conversational environments like 𝜏 ‐bench, these agents often struggle with consistent reasonin

Cited by 0SourcePDFScholar
2024

LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models

ACL 2024long

Recently developed large language models (LLMs) have been shown to perform remarkably well on a wide range of language understanding tasks. But, can they really “reason” over the natural language? This question has been receiving significant research attention and many reasoning skills such as commo…

2024

Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models

EMNLP 2024main

As Large Language Models (LLMs) continue to exhibit remarkable performance in natural language understanding tasks, there is a crucial need to measure their ability for human-like multi-step logical reasoning. Existing logical reasoning evaluation benchmarks often focus primarily on simplistic singl…

2024

Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?

EMNLP 2024main

Solving grid puzzles involves a significant amount of logical reasoning. Hence, it is a good domain to evaluate reasoning capability of a model which can then guide us to improve the reasoning ability of models. However, most existing works evaluate only the final predicted answer of a puzzle, witho…

2023

LogicAttack: Adversarial Attacks for Evaluating Logical Consistency of Natural Language Inference

EMNLP 2023short findings

Recently Large Language Models (LLMs) such as GPT-3, ChatGPT, and FLAN have led to impressive progress in Natural Language Inference (NLI) tasks. However, these models may rely on simple heuristics or artifacts in the evaluation data to achieve their high performance, which suggests that they still…

Cited by 0SourcecodeScholar