paper-with-me

Papers

Evaluation Framework for AI Systems in "the Wild"

2025-04-23 · Sarah Jabbour, Trenton Chang, Anindya Das Antar, Joseph Peper, Insu Jang, Jiachen Liu, Jae-Won Chung, Shiqi He, Michael Wellman, Bryan Goodman, Elizabeth Bondi-Kelly, Kevin Samy, Rada Mihalcea, Mosharaf Chowdhury, David Jurgens, Lu Wang

Generative AI (GenAI) models have become vital across industries, yet current evaluation methods have not adapted to their widespread use. Traditional evaluations often rely on benchmarks and fixed datasets, frequently failing to reflect real-world performance, which creates a gap between lab-tested outcomes and practical applications. This white paper proposes a comprehensive framework for how we should evaluate real-world GenAI systems, emphasizing diverse, evolving inputs and holistic, dynamic, and ongoing assessment approaches. The paper offers guidance for practitioners on how to design evaluation methods that accurately reflect real-time capabilities, and provides policymakers with recommendations for crafting GenAI policies focused on societal impacts, rather than fixed performance numbers or parameter sizes. We advocate for holistic frameworks that integrate performance, fairness, and ethics and the use of continuous, outcome-oriented methods that combine human and automated assessments while also being transparent to foster trust among stakeholders. Implementing these strategies ensures GenAI models are not only technically proficient but also ethically responsible and impactful.

📄 PDF Abstract BibTeX arXiv:2504.16778

Code (0)

등록된 구현이 없습니다.

Tasks

EthicsFairness

Similar Papers 제목 키워드 기반

Does Your Wildfire Prediction Model Actually Work, or Just Score Well?

2026-05-14 · Yangshuang Xu, Yuyang Dai, Liling Chang, Qi Wang 외 arxiv

Wildfire prediction is important for early warning and resource allocation, yet existing Earth foundation models (Earth FMs) are pretrained for general atmospheric and geophysical objectives rather than wildfire forecast…

SUDO: a framework for evaluating clinical artificial intelligence systems without ground-truth annotations

2024-01-02 · Dani Kiyasseh, Aaron Cohen, Chengsheng Jiang, Nicholas Altieri

A clinical artificial intelligence (AI) system is often validated on a held-out set of data which it has not been exposed to before (e.g., data from a different hospital with a distinct electronic health record system). …

ShadowWolf -- Automatic Labelling, Evaluation and Model Training Optimised for Camera Trap Wildlife Images

2025-12-06 · Jens Dede, Anna Förster arxiv

The continuous growth of the global human population is leading to the expansion of human habitats, resulting in decreasing wildlife spaces and increasing human-wildlife interactions. These interactions can range from mi…

WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery

2026-02-09 · Aydin Ayanzadeh, Prakhar Dixit, Sadia Kamal, Milton Halem arxiv

Wildfires are a growing threat to ecosystems, human lives, and infrastructure, with their frequency and intensity rising due to climate change and human activities. Early detection is critical, yet satellite-based monito…

MolRecBench-Wild: A Real-World Benchmark for Optical Chemical Structure Recognition

2026-05-07 · Haote Yang, Hui Wang, Chen Zhu, Jingchao Wang 외 arxiv

Optical Chemical Structure Recognition (OCSR) aims to translate molecular diagrams in scientific literature into machine-readable formats, but current systems remain unreliable on real-world images due to substantial vis…