paper-with-me

홈 › Papers

PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures

2026-05-15 · Yunan Lu, Luigi Liu, Omar Yahia, Arpit Sharma, Zhou Yu arxiv

Evaluating LLM-based agents remains challenging because identifying meaningful failure cases often requires substantial human effort to design realistic test scenarios. Prior works primarily focus on automatically discovering agent failures induced by adversarial users, while overlooking queries with real user intents that also trigger agent failures. We introduce PQR, a framework that not only surfaces agent failures with respect to specific objectives (e.g., helpfulness, safety, etc.) but also resembles real users' intents. PQR operates through an iterative interaction between two complementary modules. The query refinement module performs rewrites to explore diverse query variations, while the prompt refinement module uses prior feedback to derive new objective-violating strategies and realism policies for refining prompts, which in turn generate failure-triggering yet realistic queries. We evaluate PQR on detecting an e-commerce QA agent's unhelpful responses. Our method uncovers 23% - 78% more unhelpful responses, and our generated queries are more diverse and realistic compared to previous methods.

📄 PDF Abstract BibTeX arXiv:2605.16551

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance

2025-12-18 · Shubh Laddha, Lucas Changbencharoen, Win Kuptivej, Surya Shringla 외 arxiv

Model Context Protocol (MCP) servers contain a collection of thousands of open-source standardized tools, linking LLMs to external systems; however, existing datasets and benchmarks lack realistic, human-like user querie…

How Would You Say It? Eliciting Lexically Diverse Dialogue for Supervised Semantic Parsing

2017-08-01 · WS 2017 8 · Ravich, Abhilasha er, Thomas Manzini, Matthias Grabmair 외

Building dialogue interfaces for real-world scenarios often entails training semantic parsers starting from zero examples. How can we build datasets that better capture the variety of ways users might phrase their querie…

Semantic Parsing

SynthTRIPs: A Knowledge-Grounded Framework for Benchmark Query Generation for Personalized Tourism Recommenders

2025-04-12 · Ashmi Banerjee, Adithi Satish, Fitri Nur Aisyah, Wolfgang Wörndl 외

Tourism Recommender Systems (TRS) are crucial in personalizing travel experiences by tailoring recommendations to users' preferences, constraints, and contextual factors. However, publicly available travel datasets often…

HallucinationRecommendation Systems

SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL Workloads

2025-07-08 · Jiale Lao, Immanuel Trummer

Database research and development often require a large number of SQL queries for benchmarking purposes. However, acquiring real-world SQL queries is challenging due to privacy concerns, and existing SQL generation metho…

Benchmarking

CodeScholar: Growing Idiomatic Code Examples

2023-12-23 · Manish Shetty, Koushik Sen, Ion Stoica

Programmers often search for usage examples for API methods. A tool that could generate realistic, idiomatic, and contextual usage examples for one or more APIs would be immensely beneficial to developers. Such a tool wo…

Program Synthesis