paper-with-me

Papers

EvalAgent: Discovering Implicit Evaluation Criteria from the Web

2025-04-21 · Manya Wadhwa, Zayne Sprague, Chaitanya Malaviya, Philippe Laban, Junyi Jessy Li, Greg Durrett

Evaluation of language model outputs on structured writing tasks is typically conducted with a number of desirable criteria presented to human evaluators or large language models (LLMs). For instance, on a prompt like "Help me draft an academic talk on coffee intake vs research productivity", a model response may be evaluated for criteria like accuracy and coherence. However, high-quality responses should do more than just satisfy basic task requirements. An effective response to this query should include quintessential features of an academic talk, such as a compelling opening, clear research questions, and a takeaway. To help identify these implicit criteria, we introduce EvalAgent, a novel framework designed to automatically uncover nuanced and task-specific criteria. EvalAgent first mines expert-authored online guidance. It then uses this evidence to propose diverse, long-tail evaluation criteria that are grounded in reliable external sources. Our experiments demonstrate that the grounded criteria produced by EvalAgent are often implicit (not directly stated in the user's prompt), yet specific (high degree of lexical precision). Further, EvalAgent criteria are often not satisfied by initial responses but they are actionable, such that responses can be refined to satisfy them. Finally, we show that combining LLM-generated and EvalAgent criteria uncovers more human-valued criteria than using LLMs alone.

📄 PDF Abstract BibTeX arXiv:2504.15219

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Empirical Study of Automating Agent Evaluation

2026-05-12 · Kang Zhou, Sangmin Woo, Haibo Ding, Kiran Ramnath 외 arxiv

Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive. A natural question arises: can frontier coding assistants reliably…

The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research

2026-02-05 · Xiaoyan Bai, Alexander Baumgartner, Haojia Sun, Ari Holtzman 외 arxiv

Reproducibility crises across sciences highlight the limitations of the paper-centric review system in assessing the rigor and reproducibility of research. AI agents that autonomously design and generate large volumes of…

The Montparnasse Algorithm for RNA Design

2026-05-25 · Tristan Cazenave arxiv

RNA design consists of discovering a nucleotide sequence that optimizes predefined criteria, such as secondary structure. It is useful for synthetic biology, medicine, and nanotechnology. We propose Montparnasse, a Monte…

APD-Agents: A Large Language Model-Driven Multi-Agents Collaborative Framework for Automated Page Design

2025-11-18 · Xinpeng Chen, Xiaofeng Han, Kaihao Zhang, Guochao Ren 외 arxiv

Layout design is a crucial step in developing mobile app pages. However, crafting satisfactory designs is time-intensive for designers: they need to consider which controls and content to present on the page, and then re…

Criteria for the Annotation of Implicit Stereotypes

2022-06-01 · LREC 2022 6 · Wolfgang Schmeisser-Nieto, Montserrat Nofre, Mariona Taulé

The growth of social media has brought with it a massive channel for spreading and reinforcing stereotypes. This issue becomes critical when the affected targets are minority groups such as women, the LGBT+ community and…

Sentence