paper-with-me

홈 › Papers

YIELD: A Large-Scale Dataset and Evaluation Framework for Information Elicitation Agents

2026-04-13 · Victor De Lima, Grace Hui Yang arxiv

Most conversational agents (CAs) are designed to satisfy user needs through user-driven interactions. However, many real-world settings, such as academic interviewing, judicial proceedings, and journalistic investigations, involve broader institutional decision-making processes and require agents that can elicit information from users. In this paper, we introduce Information Elicitation Agents (IEAs) in which the agent's goal is to elicit information from users to support the agent's institutional or task-oriented objectives. To enable systematic research on this setting, we present YIELD, a 26M-token dataset of 2,281 ethically sourced, human-to-human dialogues. Moreover, we formalize information elicitation as a finite-horizon POMDP and propose novel metrics tailored to IEAs. Pilot experiments on multiple foundation LLMs show that training on YIELD improves their alignment with real elicitation behavior and findings are corroborated by human evaluation. We release YIELD under CC BY 4.0. The dataset, project code, evaluation tools, and fine-tuned model adapters are available at: https://github.com/infosenselab/yield.

📄 PDF Abstract BibTeX arXiv:2604.10968

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation

2025-10-21 · Abeer Badawi, Elahe Rahimi, Md Tahmid Rahman Laskar, Sheri Grach 외 arxiv

Evaluating Large Language Models (LLMs) for mental health support is challenging due to the emotionally and cognitively complex nature of therapeutic dialogue. Existing benchmarks are limited in scale, reliability, often…

JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Engines

2026-06-18 · Jianwen Sun, Chuanhao Li, Zizhen Li, Yukang Feng 외 arxiv

Current AI-driven game development has made substantial progress in asset generation, gameplay design, and web-based game coding, yet project-level code engineering on professional game engines remains largely unexplored…

Code Completion

SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis

2025-06-09 · Jianhui Wei, Zikai Xiao, Danyu Sun, Luqi Gong 외

Surgical video understanding is pivotal for enabling automated intraoperative decision-making, skill assessment, and postoperative quality improvement. However, progress in developing surgical video foundation models (FM…

Action ClassificationBenchmarkingDecision MakingDomain Generalization+2

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

2026-08-27 · Mingqi Gao, Anthony Sicilia, Weiyan Shi arxiv

Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered e…

HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

2024-02-06 · Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 외

Automated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a standardized evaluation framework to rigoro…

Red Teaming