paper-with-me

홈 › Papers

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability

2026-05-12 · Kyosuke Takami, Yuka Tateisi, Satoshi Sekine, Yusuke Miyao arxiv

Authentic school examinations provide a high-validity test bed for evaluating multimodal large language models (MLLMs), yet benchmarks grounded in Japanese K-12 assessments remain scarce. We present a multimodal dataset constructed from Japan's National Assessment of Academic Ability, comprising officially released middle-school items in Science, Mathematics, and Japanese Language. Unlike existing benchmarks based on synthetic or curated data, our dataset preserves real exam layouts, diagrams, and Japanese educational text, together with nationwide aggregated student response distributions (N $\approx$ 900{,}000). These features enable direct comparison between human and model performance under a unified evaluation framework. We benchmark recent multimodal LLMs using exact-match accuracy and character-level F1 for open-ended responses, observing substantial variation across subjects and strong sensitivity to visual reasoning demands. Human evaluation and LLM-as-judge analyses further assess the reliability of automatic scoring. Our dataset establishes a reproducible, human-grounded benchmark for multimodal educational reasoning and supports future research on evaluation, feedback generation, and explainable AI in authentic assessment contexts. Our dataset is available at: https://github.com/KyosukeTakami/gakucho-benchmark

📄 PDF Abstract BibTeX arXiv:2605.11663

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

2026-09-09 · Yizhan Li, Jianxin You, Mengyang Xiong, Yinhuan Chen 외 hf

Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a meaningful test of embodied intelligence and a hard requirement for deploying multimodal large language models (MLLMs) as…

Question Answering

NoRA: Evaluating Grounded Reasonableness in Visual First-person Normative Action Reasoning

2026-06-03 · Sichao Li, Sai Ma, Daniel Kilov, Secil Yanik Guyot 외 arxiv

LLMs and agentic systems are increasingly deployed in social environments, making normative competence critical for safe and appropriate behavior. However, existing approaches either assess normative judgment in text alo…

MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment

2026-02-03 · Eunkyu Park, Wesley Hanwen Deng, Cheyon Jin, Matheus Kunzler Maldaner 외 arxiv

Vision-Language Models (VLMs) continue to struggle to make morally salient judgments in multimodal and socially ambiguous contexts. Prior works typically rely on binary or pairwise supervision, which often fail to captur…

A Multimodal Framework for Human-Multi-Agent Interaction

2026-03-24 · Shaid Hasan, Breenice Lee, Sujan Sarker, Tariq Iqbal arxiv

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a u…

Multimodal Reasoning

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

2025-10-07 · Firoj Alam, Ali Ezzat Shahroor, Md. Arid Hasan, Zien Sheikh Ali 외 arxiv

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low…

Visual Question AnsweringObject Recognition