← Search

Soham Ray

2 accepted papers

2026

$\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

ICML 2026poster

Full-duplex voice agents—systems that listen and speak simultaneously—are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We introduce $\tau$-voice, a benchmark for evaluating voice agents on grounded tasks w…

Cited by 9SourceScholar
2026

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

ICML 2026oral

Existing benchmarks for conversational AI agents simulate *single-control* environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to activel…

Cited by 0SourceScholar