← Search

Kevin Zhu

4 accepted papers

2026

Reliable Weak-to-Strong Monitoring of LLM Agents

ICLR 2026oral

We stress test monitoring systems for detecting covert misbehavior in LLM agents (e.g., secretly exfiltrating data). We propose a monitor red teaming (MRT) workflow that varies agent and monitor awareness, adversarial evasion strategies, and evaluation across tool-calling (SHADE-Arena) and computer-…

Cited by 0SourcecodeScholar
2025

EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models

EMNLP 2025

The diversity of human language, shaped by social, cultural, and regional influences, presents significant challenges for natural language processing (NLP) systems. Existing benchmarks often overlook intra-language variations, leaving speakers of non-standard dialects underserved. To address this ga

Cited by 0SourcePDFScholar
2025

NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts

EMNLP 2025

Current large language models (LLMs) struggle to answer questions that span tens of thousands of tokens, especially when multi-hop reasoning is involved. While prior benchmarks explore long-context comprehension or multi-hop reasoning in isolation, none jointly vary context length and reasoning dept