τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $τ^2$-bench, with four key contributions: 1) A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2) A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3) A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4) Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, $τ^2$-bench provides a controlled testbed for agents that must both reason effectively and guide user actions.

MultiWOZ - A Large-ScaleMulti-Domain…MultiWOZ - A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue ModellingDecoupling Strategy andGeneration in…Decoupling Strategy and Generation in Negotiation Dialoguesτ-bench: A Benchmark forTool-Agent-User…τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsSWE-bench: Can LanguageModels Resolve…SWE-bench: Can Language Models Resolve Real-World GitHub Issues?WebArena: A RealisticWeb Environment for…WebArena: A Realistic Web Environment for Building Autonomous AgentsAgentBench: EvaluatingLLMs as AgentsAgentBench: Evaluating LLMs as AgentsMetaTool Benchmark forLarge Language Models…MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseIdentifying the Risks ofLM Agents with an…Identifying the Risks of LM Agents with an LM-Emulated SandboxAPIGen-MT: AgenticPipeline for Multi-Turn…APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human InterplayToolSandbox: A Stateful,Conversational…ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use CapabilitiesIntellAgent: AMulti-Agent Framework…IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI SystemsMultiAgentBench :Evaluating the…MultiAgentBench : Evaluating the Collaboration and Competition of LLM agentsAutoForge: AutomatedEnvironment Synthesis…AutoForge: Automated Environment Synthesis for Agentic Reinforcement LearningMUA-RL: Multi-turnUser-interacting Agent…MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool useDeepSeek-V3.2: Pushingthe Frontier of Open…DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsHolistic AgentLeaderboard: The Missin…Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationToolMind TechnicalReport: A Large-Scale…ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use DatasetTowards General AgenticIntelligence via…Towards General Agentic Intelligence via Environment ScalingTopoCurate:ModelingInteraction Topology fo…TopoCurate:Modeling Interaction Topology for Tool-Use Agent TrainingAgent-World: ScalingReal-World Environment…Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent IntelligenceBenchmark Test-TimeScaling of General LLM…Benchmark Test-Time Scaling of General LLM AgentsAgentNoiseBench:Benchmarking Robustness…AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy ConditionMiMo-V2-Flash TechnicalReportMiMo-V2-Flash Technical ReportScaleEnv: ScalingEnvironment Synthesis…ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Trainingτ2-Bench: EvaluatingConversational Agents i…τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。