paper-with-me

Papers

$τ$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge

2026-03-04 · Quan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan, Victor Barres arxiv

Conversational agents are increasingly deployed in knowledge-intensive settings, where correct behavior depends on retrieving and applying domain-specific knowledge from large, proprietary, and unstructured corpora during live interactions with users. Yet most existing benchmarks evaluate retrieval or tool use independently of each other, creating a gap in realistic, fully agentic evaluation over unstructured data in long-horizon interactions. We introduce $τ$-Knowledge, an extension of $τ$-Bench for evaluating agents in environments where success depends on coordinating external, natural-language knowledge with tool outputs to produce verifiable, policy-compliant state changes. Our new domain, $τ$-Banking, models realistic fintech customer support workflows in which agents must navigate roughly 700 interconnected knowledge documents while executing tool-mediated account updates. Across embedding-based retrieval and terminal-based search, even frontier models with high reasoning budgets achieve only $\sim$25.5% pass^1, with reliability degrading sharply over repeated trials. Agents struggle to retrieve the correct documents from densely interlinked knowledge bases and to reason accurately over complex internal policies. Overall, $τ$-Knowledge provides a realistic testbed for developing agents that integrate unstructured knowledge in human-facing deployments.

📄 PDF Abstract BibTeX arXiv:2603.04370

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs

2026-04-14 · Manoj Madushanka Perera, Adnan Mahmood, Kasun Eranda Wijethilake, Quan Z. Sheng arxiv

Recent advancements in Large Language Models (LLMs) have improved their ability to process extended conversational contexts, yet fine-tuning and evaluating short- and long-term memories remain difficult due to the absenc…

Knowledge Graphs

Extending Neural Generative Conversational Model using External Knowledge Sources

2018-09-14 · EMNLP 2018 10 · Prasanna Parthasarathi, Joelle Pineau

The use of connectionist approaches in conversational agents has been progressing rapidly due to the availability of large corpora. However current generative dialogue models often lack coherence and are content poor. Th…

SUQL: Conversational Search over Structured and Unstructured Data with Large Language Models

2023-11-16 · Shicheng Liu, Jialiang Xu, Wesley Tjangnaka, Sina J. Semnani 외

While most conversational agents are grounded on either free-text or structured knowledge, many knowledge corpora consist of hybrid sources. This paper presents the first conversational agent that supports the full gener…

Conversational SearchIn-Context LearningInformation Retrieval

Evaluating Agent Interactions Through Episodic Knowledge Graphs

2022-09-22 · CCGPK (COLING) 2022 10 · Selene Báez Santamaría, Piek Vossen, Thomas Baier

We present a new method based on episodic Knowledge Graphs (eKGs) for evaluating (multimodal) conversational agents in open domains. This graph is generated by interpreting raw signals during conversation and is able to …

Knowledge Graphs

On Evaluating and Comparing Open Domain Dialog Systems

2018-01-11 · Anu Venkatesh, Chandra Khatri, Ashwin Ram, Fenfei Guo 외

Conversational agents are exploding in popularity. However, much work remains in the area of non goal-oriented conversations, despite significant growth in research interest over recent years. To advance the state of the…

Goal-Oriented Dialogue SystemsOpen-Domain Dialog