← Search

Chiwei Zhu

8 accepted papers

2026

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

ICLR 2026poster

Deep Research Agents (DRAs) are emerging as one of the most practical classes of LLM-based agents. Given an open-ended research task, they find, analyze, and synthesize large numbers of online sources to produce a comprehensive report at the level of a research analyst. This can compress hours of ma…

Cited by 0SourcecodeScholar
2026

MCP-AgentBench: Evaluating Real-World Language Agent Performance with MCP-Mediated Tools

AAAI 2026technical

The Model Context Protocol (MCP) is rapidly emerging as a pivotal open standard, designed to enhance agent-tool integration and interoperability, and is positioned to unlock a new era of powerful, interconnected, and genuinely utilitarian agentic AI. However, despite MCP

Cited by 0SourcePDFScholar
2025

Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach

EMNLP 2025

Creative writing is a key capability of Large Language Models (LLMs), with potential applications in literature, storytelling, and various creative domains. However, evaluating the creativity of machine-generated texts remains a significant challenge, as existing methods either rely on costly manual

Cited by 0SourcePDFScholar
2025

From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding

ACL 2025long

The pursuit of diverse, complex, and large-scale instruction data is crucial for automatically aligning large language models (LLMs). While there are methods capable of generating synthetic instructions at scale, they either suffer from limited grounding sources, leading to a narrow distribution, or…

2025

Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

ACL 2025finding

Training language models with rationales augmentation has been shown to be beneficial in many existing works. In this paper, we identify that such a prevailing view does not hold consistently. We conduct comprehensive investigations to thoroughly inspect the impact of rationales on model performance…

2024

KNN-Instruct: Automatic Instruction Construction with K Nearest Neighbor Deduction

EMNLP 2024main

Supervised fine-tuning (SFT) is a critical procedure for aligning large language models. Despite its efficiency, the construction of SFT data often struggles with issues of quality, diversity, and scalability. Many existing methods, inspired by the Self-Instruct framework, typically generate synthet…

2023

Grammatical Error Correction via Mixed-Grained Weighted Training

EMNLP 2023long findings

The task of Grammatical Error Correction (GEC) aims to automatically correct grammatical errors in natural texts. Almost all previous works treat annotated training data equally, but inherent discrepancies in data are neglected. In this paper, the inherent discrepancies are manifested in two aspect…

Cited by 0SourceScholar
2023

On the Calibration of Large Language Models and Alignment

EMNLP 2023long findings

As large language models attract increasing attention and find widespread application, concurrent challenges of reliability also arise at the same time. Confidence calibration, an effective analysis method for gauging the reliability of deep models, serves as a crucial tool for assessing and improvi…

Cited by 0SourceScholar