paper-with-me

Papers

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

2026-05-29 · Yeil Jeong, Youngjin Yoo, Jiyoung Bae, Seobin Sohn, Hyejin Han, Jinseo Lee, Howard Scott, Unggi Lee arxiv

Classroom videos contain observable teaching practices, but their pedagogical and visual signals are rarely organized in forms suitable for model evaluation. We present \textit{TeachObs}, a human-validated benchmark for multimodal teaching observation in classroom videos. \textit{TeachObs} includes 30 public lesson videos from eight countries divided into 5,158 fixed 15-second scenes. Seven researchers annotated each scene with 39 binary observation codes, covering 20 visual codes, such as gesture, board work, pointing, and visual materials, and 19 nonvisual codes, such as instruction, monitoring, questioning, feedback, and reflection. Gold segment labels are constructed using reliability- and prevalence-aware rules based on Krippendorff's alpha. In addition to segment-level labels, three expert raters produced lesson-level ratings and qualitative evaluations of instructional design, instructional delivery, learner response, learning materials, and lesson closure across the 30 lessons, with rater coverage detailed in the body. Using these two human reference layers, we evaluate five vision-capable frontier LLMs across three tracks - text-only segment coding, text + frame segment coding, and lesson-level coverage scored under an LLM-as-judge protocol - and find that no single model consistently outperforms others across all three tracks, that adding a mid-frame inflates both true and false attributions per scene, and that model evaluations over-rate procedurally clear lessons relative to expert raters. \textit{TeachObs} therefore supports both fine-grained annotation benchmarking and whole-lesson evaluation, showing where AI systems can assist classroom video analysis and where expert judgment remains necessary across varied subjects, classroom formats, and annotation difficulty levels.

📄 PDF Abstract BibTeX arXiv:2605.30673

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering

2023-05-05 · Lei Wang, Yi Hu, Jiabang He, Xing Xu 외

Large Language Models (LLMs) have recently demonstrated exceptional performance in various Natural Language Processing (NLP) tasks. They have also shown the ability to perform chain-of-thought (CoT) reasoning to solve co…

Language ModelingLanguage ModellingLarge Language ModelQuestion Answering+1

Nonverbal Immediacy Analysis in Education: A Multimodal Computational Model

2024-07-24 · Uroš Petković, Jonas Frenkel, Olaf Hellwich, Rebecca Lazarides

This paper introduces a novel computational approach for analyzing nonverbal social behavior in educational settings. Integrating multimodal behavioral cues, including facial expressions, gesture intensity, and spatial d…

Benchmarking Local Language Models for Social Robots using Edge Devices

2026-05-04 · Dorian Lamouille, Matevž B. Zorec, Farnaz Baksh, Karl Kruusamäe arxiv

Social-educational robots designed for socially interactive pedagogical support, such as the Robot Study Companion (RSC), rely on responsive, privacy-preserving interaction despite severely limited compute. However, ther…

General Knowledge

LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching

2026-06-15 · Jaward Sesay, Yue Yu, Siwei Dong, Börje F. Karlsson arxiv

Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but also dynamically adapt their instruction to diverse learners. However, existing …

Semantic Segmentation

MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding

2025-05-26 · Jeonghun Baek, Kazuki Egashira, Shota Onohara, Atsuyuki Miyai 외

Manga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga c…

Question AnsweringVisual Question Answering