paper-with-me

홈 › Papers

ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding

2025-10-15 · Xiaozhe Li, TianYi Lyu, Siyi Yang, Yuxi Gong, Yizhao Yang, Jinxuan Huang, Ligao Zhang, Zhuoyi Huang, Qingwen Liu arxiv

Understanding human intent is a complex, high-level task for large language models (LLMs), requiring analytical reasoning, contextual interpretation, dynamic information aggregation, and decision-making under uncertainty. Real-world public discussions, such as consumer product discussions, are rarely linear or involve a single user. Instead, they are characterized by interwoven and often conflicting perspectives, divergent concerns, goals, emotional tendencies, as well as implicit assumptions and background knowledge about usage scenarios. To accurately understand such explicit public intent, an LLM must go beyond parsing individual sentences; it must integrate multi-source signals, reason over inconsistencies, and adapt to evolving discourse, similar to how experts in fields like politics, economics, or finance approach complex, uncertain environments. Despite the importance of this capability, no large-scale benchmark currently exists for evaluating LLMs on real-world human intent understanding, primarily due to the challenges of collecting real-world public discussion data and constructing a robust evaluation pipeline. To bridge this gap, we introduce \bench, the first dynamic, live evaluation benchmark specifically designed for intent understanding, particularly in the consumer domain. \bench is the largest and most diverse benchmark of its kind, supporting real-time updates while preventing data contamination through an automated curation pipeline.

📄 PDF Abstract BibTeX arXiv:2510.13499

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language

2025-05-23 · Naghmeh Jamali, Milad Mohammadi, Danial Baledi, Zahra Rezvani 외

Medical consumer question answering (CQA) is crucial for empowering patients by providing personalized and reliable health information. Despite recent advances in large language models (LLMs) for medical QA, consumer-ori…

BenchmarkingQuestion Answering

LLM-Based Multi-Agent System for Simulating and Analyzing Marketing and Consumer Behavior

2025-10-20 · Man-Lin Chu, Lucian Terhorst, Kadin Reed, Tom Ni 외 arxiv

Simulating consumer decision-making is vital for designing and evaluating marketing strategies before costly real-world deployment. However, post-event analyses and rule-based agent-based models (ABMs) struggle to captur…

Evaluating LLMs' Effectiveness on Real-World Consumer Device Repair Questions

2026-06-02 · Atm Mizanur Rahman, Md Arid Hasan, Syed Ishtiaque Ahmed, Sharifa Sultana arxiv

Consumer device repair is an important but underexplored testbed for large language models (LLMs). Repair tasks require reasoning over incomplete problem descriptions, hardware-specific diagnostics, actionable troublesho…

Building Evaluation Datasets for Consumer-Oriented Information Retrieval

2016-05-01 · LREC 2016 5 · Lorraine Goeuriot, Liadh Kelly, Guido Zuccon, Joao Palotti

Common people often experience difficulties in accessing relevant, correct, accurate and understandable health information online. Developing search techniques that aid these information needs is challenging. In this pap…

Information RetrievalRetrieval

A Benchmark for Long-Form Medical Question Answering

2024-11-14 · Pedram Hosseini, Jessica M. Sin, Bing Ren, Bryceton G. Thomas 외

There is a lack of benchmarks for evaluating large language models (LLMs) in long-form medical question answering (QA). Most existing medical QA evaluation benchmarks focus on automatic metrics and multiple-choice questi…

Answer GenerationFormMedical Question AnsweringMultiple-choice+1