← Search

Sangwu Park

3 accepted papers

2026

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

ICML 2026poster

Driven by recent advancements in tool-augmented Large Language Model (LLM) agents, comprehensive benchmark datasets for evaluating these tool-augmented agents are being actively developed. Although these benchmarks incorporate increasingly complex user requests and a diverse array of tools, the eval…

Cited by 0SourceScholar
2025

Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models

EMNLP 2025

As the use of large language model (LLM) agents continues to grow, their safety vulnerabilities have become increasingly evident. Extensive benchmarks evaluate various aspects of LLM safety by defining the safety relying heavily on general standards, overlooking user-specific standards. However, saf

2025

SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials

NAACL 2025findings

Recently, interpreting complex charts with logical reasoning has emerged as challenges due to the development of vision-language models. A prior state-of-the-art (SOTA) model has presented an end-to-end method that leverages the vision-language model to convert charts into table format utilizing Lar…