Papers
arxiv:2406.12045

τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Published on Jun 17, 2024
· Submitted by
Soham Parikh
on Jun 21, 2024
Authors:
,
,

Abstract

Tau-bench tests agents' interaction with simulated users and adherence to domain-specific rules, revealing inconsistencies in function calling agents.

Existing benchmarks do not test language agents on their interaction with human users or ability to follow domain-specific rules, both of which are vital for deploying them in real world applications. We propose tau-bench, a benchmark emulating dynamic conversations between a user (simulated by language models) and a language agent provided with domain-specific API tools and policy guidelines. We employ an efficient and faithful evaluation process that compares the database state at the end of a conversation with the annotated goal state. We also propose a new metric (pass^k) to evaluate the reliability of agent behavior over multiple trials. Our experiments show that even state-of-the-art function calling agents (like gpt-4o) succeed on <50% of the tasks, and are quite inconsistent (pass^8 <25% in retail). Our findings point to the need for methods that can improve the ability of agents to act consistently and follow rules reliably.

Community

•
This comment has been hidden (marked as Off-Topic)

Is there a tau-bench leader board available? I would love to see if the newer models are improving on this "reliability" benchmark as time passes

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2406.12045
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 3

Datasets citing this paper 5

Browse 5 datasets citing this paper

Spaces citing this paper 3

Collections including this paper 8