paper-with-me

홈 › Papers

Classification of Important Segments in Educational Videos using Multimodal Features

2020-10-26 · Junaid Ahmed Ghauri, Sherzod Hakimov, Ralph Ewerth

Videos are a commonly-used type of content in learning during Web search. Many e-learning platforms provide quality content, but sometimes educational videos are long and cover many topics. Humans are good in extracting important sections from videos, but it remains a significant challenge for computers. In this paper, we address the problem of assigning importance scores to video segments, that is how much information they contain with respect to the overall topic of an educational video. We present an annotation tool and a new dataset of annotated educational videos collected from popular online learning platforms. Moreover, we propose a multimodal neural architecture that utilizes state-of-the-art audio, visual and textual features. Our experiments investigate the impact of visual and temporal information, as well as the combination of multimodal features on importance prediction.

📄 PDF Abstract BibTeX arXiv:2010.13626

Code (1)

VideoAnalysis/EDUVSUM 공식 구현 pytorch

Tasks

ClassificationGeneral Classification

Similar Papers 제목 키워드 기반

Multimodal Fusion and Coherence Modeling for Video Topic Segmentation

2024-08-01 · Hai Yu, Chong Deng, Qinglin Zhang, Jiaqing Liu 외

The video topic segmentation (VTS) task segments videos into intelligible, non-overlapping topics, facilitating efficient comprehension of video content and quick access to specific content. VTS is also critical to vario…

Contrastive LearningMixture-of-ExpertsScene SegmentationSegmentation+1

Nonverbal Immediacy Analysis in Education: A Multimodal Computational Model

2024-07-24 · Uroš Petković, Jonas Frenkel, Olaf Hellwich, Rebecca Lazarides

This paper introduces a novel computational approach for analyzing nonverbal social behavior in educational settings. Integrating multimodal behavioral cues, including facial expressions, gesture intensity, and spatial d…

Scalable and Explainable Learner-Video Interaction Prediction using Multimodal Large Language Models

2026-04-06 · Dominik Glandorf, Fares Fawzi, Tanja Käser arxiv

Learners' use of video controls in educational videos provides implicit signals of cognitive processing and instructional design quality, yet the lack of scalable and explainable predictive models limits instructors' abi…

ConfusionBench: An Expert-Validated Benchmark for Confusion Recognition and Localization in Educational Videos

2026-03-18 · Lu Dong, Xiao Wang, Mark Frank, Srirangaraj Setlur 외 arxiv

Recognizing and localizing student confusion from video is an important yet challenging problem in educational AI. Existing confusion datasets suffer from noisy labels, coarse temporal annotations, and limited expert val…

Revealing Temporal Label Noise in Multimodal Hateful Video Classification

2025-08-06 · Shuonan Yang, Tailin Chen, Rahul Singh, Jiangbei Yue 외 arxiv

The rapid proliferation of online multimedia content has intensified the spread of hate speech, presenting critical societal and regulatory challenges. While recent work has advanced multimodal hateful video detection, m…

Video Classification