paper-with-me

Papers

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

2025-04-20 · Enxin Song, Wenhao Chai, Weili Xu, Jianwen Xie, Yuxuan Liu, Gaoang Wang

Recent advancements in language multimodal models (LMMs) for video have demonstrated their potential for understanding video content, yet the task of comprehending multi-discipline lectures remains largely unexplored. We introduce Video-MMLU, a massive benchmark designed to evaluate the capabilities of LMMs in understanding Multi-Discipline Lectures. We evaluate over 90 open-source and proprietary models, ranging from 0.5B to 40B parameters. Our results highlight the limitations of current models in addressing the cognitive challenges presented by these lectures, especially in tasks requiring both perception and reasoning. Additionally, we explore how the number of visual tokens and the large language models influence performance, offering insights into the interplay between multimodal perception and reasoning in lecture comprehension.

📄 PDF Abstract BibTeX arXiv:2504.14693

Code (1)

espere-1119-song/video-mmlu 공식 구현 pytorch

Tasks

MMLU

Similar Papers 제목 키워드 기반

Efficient Indexing of Meta-Data (Extracted from Educational Videos)

2023-12-11 · Shalika Kumbham, Abhijit Debnath, Krothapalli Sreenivasa Rao

Video lectures are becoming more popular and in demand as online classroom teaching is becoming more prevalent. Massive Open Online Courses (MOOCs), such as NPTEL, have been creating high-quality educational content that…

A Deep Dive into the Disparity of Word Error Rates Across Thousands of NPTEL MOOC Videos

2023-07-20 · Anand Kumar Rai, Siddharth D Jaiswal, Animesh Mukherjee

Automatic speech recognition (ASR) systems are designed to transcribe spoken language into written text and find utility in a variety of applications including voice assistants and transcription services. However, it has…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

The OpenCourseWare Metadiscourse (OCWMD) Corpus

2016-05-01 · LREC 2016 5 · Ghada Alharbi, Thomas Hain

This study describes a new corpus of over 60,000 hand-annotated metadiscourse acts from 106 OpenCourseWare lectures, from two different disciplines: Physics and Economics. Metadiscourse is a set of linguistic expressions…

Your click decides your fate: Inferring Information Processing and Attrition Behavior from MOOC Video Clickstream Interactions

2014-07-26 · WS 2014 10 · Tanmay Sinha, Patrick Jermann, Nan Li, Pierre Dillenbourg

In this work, we explore video lecture interaction in Massive Open Online Courses (MOOCs), which is central to student learning experience on these educational platforms. As a research contribution, we operationalize vid…

LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?

2026-01-27 · Zhuang Yu, Lei Shen, Jing Zhao, Shiliang Sun arxiv

Recent multimodal large language models (MLLMs) have shown remarkable progress across vision, audio, and language tasks, yet their performance on long-form, knowledge-intensive, and temporally structured educational cont…