paper-with-me

홈 › Papers

EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems

2026-07-06 · Kenneth Benavides, Josh Fleischer, Danti Chen arxiv

Teams deploying large language models in business contexts need evaluation systems, yet most treat evaluation as static model selection: run benchmarks, rank models, deploy the winner. This framing misses evaluation's primary value for production systems--diagnosing why a system underperforms and guiding what to fix. We present EvalLoop, a methodology for evaluation-driven iterative improvement. EvalLoop organizes evaluation around three mechanisms: (1) dimensional metric grouping that decomposes quality into business-relevant dimensions enabling orthogonal failure diagnosis; (2) failure mode classification that categorizes why outputs fail within weak dimensions, bridging diagnosis to action; and (3) a structured iteration workflow where each evaluation run varies one system variable and compares dimensional profiles before and after. We validate EvalLoop through a case study on sales intelligence briefing generation (10 models, 3 providers, 18 metrics, 5 dimensions, 3 iterations). Dimensional diagnosis identified that 69% of hallucination failures were prompt-induced interpretation errors--invisible in aggregate scoring. A targeted prompt fix improved the best model from 82.6% to 94.6% overall, with improvement concentrated in diagnosed dimensions (Content Accuracy +16.8pp, Synthesis Power +26.4pp). An undirected configuration change in a prior iteration produced zero impact, illustrating the cost of iterating without diagnosis. We additionally demonstrate that dimensional profiling enables deployment-specific model selection, and that a one-time blind human gate on a finalist panel (4 models, 16 cases) confirms dimensional rankings while resolving multi-criteria deployment trade-offs--a 94% reduction in review burden compared to evaluating the full design. EvalLoop is packaged as reusable artifacts (playbook, agent specification, template repository) for adoption by other teams.

📄 PDF Abstract BibTeX arXiv:2607.05638

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research Insights

2025-05-06 · Chengzhang Yu, Yiming Zhang, Zhixin Liu, Zenghui Ding 외

The automation of scientific research through large language models (LLMs) presents significant opportunities but faces critical challenges in knowledge synthesis and quality assurance. We introduce Feedback-Refined Agen…

Paper generation

Hypothesis-Driven Deep Research with Large Language Models: A Structured Methodology for Automated Knowledge Discovery

2026-05-11 · Michael Chin arxiv

Current AI-powered research systems adopt a direct search-then-summarize paradigm that treats hypotheses as end products of scientific discovery. We argue this leaves a critical gap: hypotheses can serve a far more power…

Information Retrieval

GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization

2025-06-25 · Martin Andrews, Sam Witteveen

Optimizing GPU kernels for high performance is a complex task, often demanding deep architectural knowledge, extensive profiling, and iterative experimentation. This challenge is amplified when targeting newer or less-do…

GPU

RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques

2025-01-24 · Zhengyang Tang, Ziniu Li, Zhenyang Xiao, Tian Ding 외

Critiques are important for enhancing the performance of Large Language Models (LLMs), enabling both self-improvement and constructive feedback for others by identifying flaws and suggesting improvements. However, evalua…

Language ModelingLanguage Modelling

A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness

2025-10-23 · Brent Winslow, Jacqueline Shreibati, Javier Perez, Hao-Wei Su 외 arxiv

The incorporation of generative artificial intelligence into personal health applications presents a transformative opportunity for personalized, data-driven health and fitness guidance, yet also poses challenges related…