paper-with-me

홈 › Papers

An Approach to Grounding AI Model Evaluations in Human-derived Criteria

2025-09-04 · Sasha Mitts arxiv

In the rapidly evolving field of artificial intelligence (AI), traditional benchmarks can fall short in attempting to capture the nuanced capabilities of AI models. We focus on the case of physical world modeling and propose a novel approach to augment existing benchmarks with human-derived evaluation criteria, aiming to enhance the interpretability and applicability of model behaviors. Grounding our study in the Perception Test and OpenEQA benchmarks, we conducted in-depth interviews and large-scale surveys to identify key cognitive skills, such as Prioritization, Memorizing, Discerning, and Contextualizing, that are critical for both AI and human reasoning. Our findings reveal that participants perceive AI as lacking in interpretive and empathetic skills yet hold high expectations for AI performance. By integrating insights from our findings into benchmark design, we offer a framework for developing more human-aligned means of defining and measuring progress. This work underscores the importance of user-centered evaluation in AI development, providing actionable guidelines for researchers and practitioners aiming to align AI capabilities with human cognitive processes. Our approach both enhances current benchmarking practices and sets the stage for future advancements in AI model evaluation.

📄 PDF Abstract BibTeX arXiv:2509.04676

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment

2026-08-01 · Garv Vikram Gursahaney, Baskhad Idrisov, Thorsten Fröhlich, Tim Schlippe arxiv

Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical …

LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks

2025-12-12 · Shogo Fujita, Yuji Naraki, Yiqing Zhu, Shinsuke Mori arxiv

This paper introduces LegalRikai: Open Benchmark, a new benchmark comprising four complex tasks that emulate Japanese corporate legal practices. The benchmark was created by legal professionals under the supervision of a…

Grounded but Misleading: Evaluating Semantic Alignment in AI-Generated Security Explanations

2026-02-04 · Heajun An, Connor Ng, Sandesh Sharma Dulal, Junghwan Kim 외 arxiv

Online scams increasingly leverage fluent and context-aware social engineering strategies, creating growing demand for AI systems that explain why a message may be risky. However, explanations that cite detector-derived …

Bonsai: Interpretable Tree-Adaptive Grounded Reasoning

2025-04-04 · Kate Sanders, Benjamin Van Durme

To develop general-purpose collaborative agents, humans need reliable AI systems that can (1) adapt to new domains and (2) transparently reason with uncertainty to allow for verification and correction. Black-box models …

Question AnsweringSpecificity

From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

2026-07-04 · Xiangyu Shi, Ruoxi Yang, Wei Tao, Jiwen Zhang 외 arxiv

Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation protocols mainly reward stopping within a…

Visual GroundingEntity Alignment