← Search

Victor Barres

4 accepted papers

2026

$\tau$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge

ICML 2026poster

Conversational agents are increasingly deployed in knowledge-intensive settings, where correct behavior depends on acquiring and applying domain-specific knowledge from large, proprietary, and unstructured corpora during live interactions with users. Yet most existing benchmarks evaluate retrieval o…

Cited by 0SourceScholar
2026

$\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

ICML 2026poster

Full-duplex voice agents—systems that listen and speak simultaneously—are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We introduce $\tau$-voice, a benchmark for evaluating voice agents on grounded tasks w…

Cited by 9SourceScholar
2026

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

ICML 2026oral

Existing benchmarks for conversational AI agents simulate *single-control* environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to activel…

Cited by 0SourceScholar
2025

From Generating Answers to Building Explanations: Integrating Multi-Round RAG and Causal Modeling for Scientific QA

NAACL 2025industry

Application of LLMs for complex causal question answering can be stymied by their opacity and propensity for hallucination. Although recent approaches such as Retrieval Augmented Generation and Chain of Thought prompting have improved reliability, we argue current approaches are insufficient and fur…

Cited by 0SourcePDFScholar