← Search

Pedram Razavi

2 accepted papers

2026

$\tau$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge

ICML 2026poster

Conversational agents are increasingly deployed in knowledge-intensive settings, where correct behavior depends on acquiring and applying domain-specific knowledge from large, proprietary, and unstructured corpora during live interactions with users. Yet most existing benchmarks evaluate retrieval o…

Cited by 0SourceScholar
2025

{$\tau$}-bench: A Benchmark for \underline{T}ool-\underline{A}gent-\underline{U}ser Interaction in Real-World Domains

ICLR 2025poster

Existing benchmarks for language agents do not set them up to interact with human users or follow domain-specific rules, both of which are vital to safe and realistic deployment. We propose $\tau$-bench, a benchmark with two domains (retail and airline) emulating dynamic conversations between a user…

Cited by 2SourcePDFScholar