paper-with-me

Papers

Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework

2026-06-07 · Aman Gupta, Kevin Rossell, Edesio Alcobaça, Jose Chrystian Lima Pacheco, Carolina Baptista de Lima, Shao Tang, Luiz Paulo Rabachini, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath arxiv

The rapid rise in LLM capabilities has made AI agents increasingly viable across a broad range of tasks. Among the most promising applications is building production-ready customer-facing agents, a challenge that demands coordinated excellence in evaluation methodology, context engineering, training, and online measurement. Yet these critical pillars are typically developed in isolation, creating blind spots that only surface after deployment. In this paper, we present a unified framework that bridges offline development with online impact for customer support AI agents at Nubank, a company with 100M+ users. Our approach integrates several key components: (1) structured context engineering tailored to customer support agents, (2) systematic human-in-the-loop prompt iteration, (3) rigorous LLM judge evaluation with measured inter-rater agreement and GEPA optimization for consistency, and (4) ideation-to-production validation. A central insight is that evaluation-pipeline quality directly determines iteration velocity. We present results from five production deployments spanning distinct domains: card delivery, debt management, credit-limit support, card management, and product explanation. These deployments deliver consistent customer-satisfaction gains while substantially accelerating iteration. In our card-delivery deployment, large-scale A/B testing yields a 37 percentage-point improvement in AI transactional Net Promoter Score and a 29 percentage-point gain in self-service rate over prior agent variants, alongside a strong correlation between offline simulation metrics and online outcomes, demonstrating that eval-driven development reliably predicts production impact. On most use cases, AI satisfaction reaches within a few percentage points of expert human agents.

📄 PDF Abstract BibTeX arXiv:2606.08867

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents

2025-03-05 · Jingying Zeng, Hui Liu, Zhenwei Dai, Xianfeng Tang 외

With the advancement of conversational large language models (LLMs), several LLM-based Conversational Shopping Agents (CSA) have been developed to help customers smooth their online shopping. The primary objective in bui…

AttributeIn-Context LearningMisinformation

Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems

2021-04-01 · NAACL 2021 4 · Derek Chen, Howard Chen, Yi Yang, Alex Lin 외

Existing goal-oriented dialogue datasets focus mainly on identifying slots and values. However, customer support interactions in reality often involve agents following multi-step procedures derived from explicitly-define…

Task-Oriented Dialogue Systems

Conversational Document Prediction to Assist Customer Care Agents

2020-10-05 · EMNLP 2020 11 · Jatin Ganhotra, Haggai Roitman, Doron Cohen, Nathaniel Mills 외

A frequent pattern in customer care conversations is the agents responding with appropriate webpage URLs that address users' needs. We study the task of predicting the documents that customer care agents can use to facil…

Information RetrievalPredictionRetrieval

TWEETSUMM - A Dialog Summarization Dataset for Customer Service

2021-11-01 · Findings (EMNLP) 2021 11 · Guy Feigenblat, Chulaka Gunasekara, Benjamin Sznajder, Sachindra Joshi 외

In a typical customer service chat scenario, customers contact a support center to ask for help or raise complaints, and human agents try to solve the issues. In most cases, at the end of the conversation, agents are ask…

Extractive SummarizationUnsupervised Extractive Summarization

TWEETSUMM -- A Dialog Summarization Dataset for Customer Service

2021-11-23 · Guy Feigenblat, Chulaka Gunasekara, Benjamin Sznajder, Sachindra Joshi 외

In a typical customer service chat scenario, customers contact a support center to ask for help or raise complaints, and human agents try to solve the issues. In most cases, at the end of the conversation, agents are ask…

Extractive SummarizationUnsupervised Extractive Summarization