← Search

Xinkai Yu

1 accepted papers

2026

AutoLibra: Agent Metric Induction from Open-Ended Human Feedback

ICLR 2026poster

Agents are predominantly evaluated and optimized via task success metrics, which are coarse, rely on manual design from experts, and fail to reward intermediate emergent behaviors. We propose AutoLibra, a framework for agent evaluation, that transforms open-ended human feedback e.g. “If you find tha…

Cited by 0SourcecodeScholar