paper-with-me

홈 › Papers

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

2025-05-26 · Atsunori Moteki, Shoichi Masui, Fan Yang, Yueqi Song, Yonatan Bisk, Graham Neubig, Ikuo Kusajima, Yasuto Watanabe, Hiroyuki Ishida, Jun Takahashi, Shan Jiang

This paper proposes FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agentic AI, they are required to monitor and report safety and health incidents, as well as manufacturing-related incidents, that may occur in real-world work environments. Existing agentic AI benchmarks have been limited to evaluating web tasks and are insufficient for evaluating agents in real-world work environments, where complexity increases significantly. In this paper, we define a new action space that agentic AI should possess for real world work environment benchmarks and improve the evaluation function from previous methods to assess the performance of agentic AI in diverse real-world tasks. The dataset consists of videos captured on-site and documents actually used in factories and warehouses, and tasks were created based on interviews with on-site workers and managers. Evaluation results confirmed that performance evaluation considering the characteristics of Multimodal LLM (MLLM) such as GPT-4o is feasible. Additionally, the effectiveness and limitations of the proposed new evaluation method were identified. The complete dataset (HuggingFace) and evaluation program (GitHub) can be downloaded from the following website: https://en-documents.research.global.fujitsu.com/fieldworkarena/.

📄 PDF Abstract BibTeX arXiv:2505.19662

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

2026-04-13 · Arun Sharma arxiv

We introduce compute-grounded reasoning (CGR), a design paradigm for spatial-aware research agents in which every answerable sub-problem is resolved by deterministic computation before a language model is asked to genera…

Spatial ReasoningCode Generation

A Survey on Agentic Multimodal Large Language Models

2025-10-13 · Huanjin Yao, Ruifei Zhang, Jiaxing Huang, Jingyi Zhang 외 arxiv

With the recent emergence of revolutionary autonomous agentic systems, research community is witnessing a significant shift from traditional static, passive, and domain-specific AI agents toward more dynamic, proactive, …

DexHoldem: Playing Texas Hold'em with Dexterous Embodied System

2026-05-18 · Feng Chen, Tianzhe Chu, Li Sun, Pei Zhou 외 arxiv

Evaluating embodied systems on real dexterous hardware requires more than isolated primitive skills: an agent must perceive a changing tabletop scene, choose a context-appropriate action, execute it with a dexterous hand…

Decision Making

Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations

2025-11-11 · JV Roig arxiv

Enterprise adoption of agentic AI systems requires reliable evaluation methods that reflect real-world deployment scenarios. Traditional LLM benchmarks suffer from training data contamination and fail to measure agentic …

Establishing Best Practices for Building Rigorous Agentic Benchmarks

2025-07-03 · Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun, Andy Zhang 외 arxiv

Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world ta…