paper-with-me

홈 › Papers

Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence

2026-01-02 · Sumanth Balaji, Piyush Mishra, Aashraya Sachdeva, Suraj Agrawal arxiv

Traditional customer support systems, such as Interactive Voice Response (IVR), rely on rigid scripts and lack the flexibility required for handling complex, policy-driven tasks. While large language model (LLM) agents offer a promising alternative, evaluating their ability to act in accordance with business rules and real-world support workflows remains an open challenge. Existing benchmarks primarily focus on tool usage or task completion, overlooking an agent's capacity to adhere to multi-step policies, navigate task dependencies, and remain robust to unpredictable user or environment behavior. In this work, we introduce JourneyBench, a benchmark designed to assess policy-aware agents in customer support. JourneyBench leverages graph representations to generate diverse, realistic support scenarios and proposes the User Journey Coverage Score, a novel metric to measure policy adherence. We evaluate multiple state-of-the-art LLMs using two agent designs: a Static-Prompt Agent (SPA) and a Dynamic-Prompt Agent (DPA) that explicitly models policy control. Across 703 conversations in three domains, we show that DPA significantly boosts policy adherence, even allowing smaller models like GPT-4o-mini to outperform more capable ones like GPT-4o. Our findings demonstrate the importance of structured orchestration and establish JourneyBench as a critical resource to advance AI-driven customer support beyond IVR-era limitations.

📄 PDF Abstract BibTeX arXiv:2601.00596

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ReviewSense: Transforming Customer Review Dynamics into Actionable Business Insights

2025-10-18 · Siddhartha Krothapalli, Kartikey Singh Bhandari, Tridib Kumar Das, Praveen Kumar 외 arxiv

As customer feedback becomes increasingly central to strategic growth, the ability to derive actionable insights from unstructured reviews is essential. While traditional AI-driven systems excel at predicting user prefer…

Sentiment Analysis

AdaCoach: A Virtual Coach for Training Customer Service Agents

2022-04-27 · Shuang Peng, Shuai Zhu, Minghui Yang, Haozhou Huang 외

With the development of online business, customer service agents gradually play a crucial role as an interface between the companies and their customers. Most companies spend a lot of time and effort on hiring and traini…

Dialogue Evaluation

Utilisation of open intent recognition models for customer support intent detection

2023-07-31 · Rasheed Mohammad, Oliver Favell, Shariq Shah, Emmett Cooper 외

Businesses have sought out new solutions to provide support and improve customer satisfaction as more products and services have become interconnected digitally. There is an inherent need for businesses to provide or out…

Intent DetectionIntent DiscoveryIntent RecognitionTransfer Learning

CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions

2025-05-24 · Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal 외

While AI agents hold transformative potential in business, effective performance benchmarking is hindered by the scarcity of public, realistic business data on widely used platforms. Existing benchmarks often lack fideli…

Benchmarking

Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews

2026-01-17 · Kartikey Singh Bhandari, Tanish Jain, Archit Agrawal, Dhruv Kumar 외 arxiv

Customer reviews contain valuable signals about service quality, but converting large-scale review corpora into actionable business recommendations remains difficult. Standard sentiment/aspect analysis is largely descrip…