paper-with-me

홈 › Papers

CXMArena: Unified Dataset to benchmark performance in realistic CXM Scenarios

2025-05-14 · Raghav Garg, Kapil Sharma, Karan Gupta

Large Language Models (LLMs) hold immense potential for revolutionizing Customer Experience Management (CXM), particularly in contact center operations. However, evaluating their practical utility in complex operational environments is hindered by data scarcity (due to privacy concerns) and the limitations of current benchmarks. Existing benchmarks often lack realism, failing to incorporate deep knowledge base (KB) integration, real-world noise, or critical operational tasks beyond conversational fluency. To bridge this gap, we introduce CXMArena, a novel, large-scale synthetic benchmark dataset specifically designed for evaluating AI in operational CXM contexts. Given the diversity in possible contact center features, we have developed a scalable LLM-powered pipeline that simulates the brand's CXM entities that form the foundation of our datasets-such as knowledge articles including product specifications, issue taxonomies, and contact center conversations. The entities closely represent real-world distribution because of controlled noise injection (informed by domain experts) and rigorous automated validation. Building on this, we release CXMArena, which provides dedicated benchmarks targeting five important operational tasks: Knowledge Base Refinement, Intent Prediction, Agent Quality Adherence, Article Search, and Multi-turn RAG with Integrated Tools. Our baseline experiments underscore the benchmark's difficulty: even state of the art embedding and generation models achieve only 68% accuracy on article search, while standard embedding methods yield a low F1 score of 0.3 for knowledge base refinement, highlighting significant challenges for current models necessitating complex pipelines and solutions over conventional techniques.

📄 PDF Abstract BibTeX arXiv:2505.09436

Code (1)

kapilsprinklr/cxmarena 공식 구현

Tasks

ArticlesRAG

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

READoc: A Unified Benchmark for Realistic Document Structured Extraction

2024-09-08 · Zichao Li, Aizier Abulaiti, Yaojie Lu, Xuanang Chen 외

Document Structured Extraction (DSE) aims to extract structured content from raw documents. Despite the emergence of numerous DSE systems, their unified evaluation remains inadequate, significantly hindering the field's …

UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation

2024-07-29 · Chaoqun Du, Yulin Wang, Jiayi Guo, Yizeng Han 외

Test-Time Adaptation (TTA) aims to adapt pre-trained models to the target domain during testing. In reality, this adaptability can be influenced by multiple factors. Researchers have identified various challenging scenar…

Test-time Adaptation

WorldScore: A Unified Evaluation Benchmark for World Generation

2025-04-01 · Haoyi Duan, Hong-Xing Yu, Sirui Chen, Li Fei-Fei 외

We introduce the WorldScore benchmark, the first unified benchmark for world generation. We decompose world generation into a sequence of next-scene generation tasks with explicit camera trajectory-based layout specifica…

Scene GenerationVideo Generation

GenShield: Unified Detection and Artifact Correction for AI-Generated Images

2026-05-15 · Zhipei Xu, Xuanyu Zhang, Youmin Xu, Qing Huang 외 arxiv

Diffusion-based image synthesis has made AI-generated images (AIGI) increasingly photorealistic, raising urgent concerns about authenticity in applications such as misinformation detection, digital forensics, and content…

FedScale: Benchmarking Model and System Performance of Federated Learning at Scale

2021-05-24 · Fan Lai, Yinwei Dai, Sanjay S. Singapuram, Jiachen Liu 외

We present FedScale, a federated learning (FL) benchmarking suite with realistic datasets and a scalable runtime to enable reproducible FL research. FedScale datasets encompass a wide range of critical FL tasks, ranging …

BenchmarkingFederated Learningimage-classificationImage Classification+6