paper-with-me

Papers

UXBench: Benchmarking User Experience in AI Assistants

2026-06-08 · Mengze Hong, Xia Zeng, Zeyang Lei, Sheng Wang, Chen Jason Zhang, Di Jiang, Taiming Fu, Jinfeng Huang, Mengqiao Liu, Qinghe Chang, Haosheng Zou, Qiongyi Zhou, Sijun He, Simonjmdeng, Haojing Huang, Zijian Li, Lucas Mu Li, Fubao Zhang, Mona Zhou, Wei Ma, Chenxuan Ma, Yuanmeng Zhang, Jian Song, Minlong Peng, Di Liang, Davey Chen arxiv

As AI assistants serve millions of users daily, evaluating user experience (UX) beyond general model capability has become increasingly important. We present UXBench, the first user-centric benchmark grounded in real user feedback signals for evaluating preference alignment and dialogue generation. The benchmark consists of three interconnected tasks, UX Judge, UX Eval, and UX Recovery, with 7,400 test instances extracted from over 70K interaction logs of a mainstream Chinese AI assistant. The dataset closely reflects real user distributions, covering 8 scenarios, 83 domains, and diverse failure patterns that pose severe challenges. Extensive experiments on 26 frontier language models provide novel insights into how well models perceive user experience and how improvements in model capability contribute to better dialogue engagement. Through comprehensive analysis of model behavior and performance gaps, we show that user feedback prediction is a learnable capability, where a reward model trained from in-the-wild feedback signals can achieve well-calibrated accuracy. We further document the systematic biases of LLM-as-a-judge evaluation protocols and compare typical response strategies that directly affect user experience. UXBench establishes a new evaluation landscape and calls for greater attention to tailored UX optimization, contributing to a user-centric scaling law that shapes the success of AI assistants.

📄 PDF Abstract BibTeX arXiv:2606.09570

Code (0)

등록된 구현이 없습니다.

Tasks

Dialogue Generation

Similar Papers 제목 키워드 기반

Reasoning for Mobile User Experience with Multimodal LLMs: Task, Benchmark, and Approach

2026-06-11 · Ruichao Mao, Zhou Fang, Teng Guo, Hao Yang 외 arxiv

User experience (UX) centered on usability, perceived consistency, and functional clarity is fundamental to real-world user interfaces (UI). The application of multimodal large language models (MLLMs) in the field of use…

Reinforcement LearningLogical ReasoningCode Generation

VoiceBench: Benchmarking LLM-Based Voice Assistants

2024-10-22 · Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao 외

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingGeneral Knowledge+2

Bilingual by default: Voice Assistants and the role of code-switching in creating a bilingual user experience

2022-06-20 · Helin Cihan, Yunhan Wu, Paola Peña, Justin Edwards 외

Conversational User Interfaces such as Voice Assistants are hugely popular. Yet they are designed to be monolingual by default, lacking support for, or sensitivity to, the bilingual dialogue experience. In this provocati…

Sensitivity

How Does Personalized Memory Shape LLM Behavior? Benchmarking Rational Preference Utilization in Personalized Assistants

2026-01-23 · Xueyang Feng, Weinan Gan, Xu Chen, Quanyu Dai 외 arxiv

Large language model (LLM)-powered assistants have recently integrated memory mechanisms that record user preferences, leading to more personalized and user-aligned responses. However, irrelevant personalized memories ar…

UXBench: Measuring the Actionability of LLM-Generated UX Critiques

2026-06-15 · Wenjie Wang, Yue Huang, Zipeng Ling, Han Bao 외 arxiv

Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs. Yet no controlled benchmark measures whether the resulting critiques are reli…