← Search

Yuzhen Huang

7 accepted papers

2026

LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth

ICML 2026poster

Frontier large language models (LLMs) are increasingly capable of carrying out long-running, real-world tasks. However, as the amount of context grows, their reliability often deteriorate, a phenomenon known as "context rot". Existing long-context benchmarks primarily focus on single-step settings t…

Cited by 0SourceScholar
2026

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

ICLR 2026poster

Large Reasoning Models (LRMs) have shown remarkable capabilities in solving complex problems through reinforcement learning (RL), particularly by generating long reasoning traces. However, these extended outputs often exhibit substantial redundancy, which limits the efficiency of LRMs. In this paper…

Cited by 0SourcecodeScholar
2026

SWE-RM: Execution-free Feedback for Software Engineering Agents

ICLR 2026poster

Execution-based feedback like unit testing is widely used in the development of coding agents through test-time scaling (TTS) and reinforcement learning (RL). This paradigm requires scalable and reliable collection of unit test cases to provide accurate feedback, and the resulting feedback is often…

Cited by 0SourceScholar
2026

The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

ICLR 2026poster

Real-world language agents must handle complex, multi-step workflows across diverse applications. For instance, an agent may manage emails by coordinating with calendars and file systems, or monitor a production database like BigQuery to detect anomalies and generate reports following a standard ope…

Cited by 0SourcecodeScholar
2025

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

ICLR 2025poster

In the absence of extensive human-annotated data for complex reasoning tasks, self-improvement -- where models are trained on their own outputs -- has emerged as a primary method for enhancing performance. Recently, the approach to self-improvement has shifted toward a more dynamic, online fashion t…

2025

Predictive Data Selection: The Data That Predicts Is the Data That Teaches

ICML 2025poster

Language model pretraining involves training on extensive corpora, where data quality plays a pivotal role. In this work, we aim to directly estimate the contribution of data during pretraining and select pretraining data in an efficient manner. Specifically, we draw inspiration from recent findings…

2023

C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

NeurIPS 2023poster

New NLP benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present C-Eval, the first comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of foundation models in a Chinese context. C-Eval comprises mu…