paper-with-me

Papers

Lecture Video Visual Objects (LVVO) Dataset: A Benchmark for Visual Object Detection in Educational Videos

2025-06-16 · Dipayan Biswas, Shishir Shah, Jaspal Subhlok

We introduce the Lecture Video Visual Objects (LVVO) dataset, a new benchmark for visual object detection in educational video content. The dataset consists of 4,000 frames extracted from 245 lecture videos spanning biology, computer science, and geosciences. A subset of 1,000 frames, referred to as LVVO_1k, has been manually annotated with bounding boxes for four visual categories: Table, Chart-Graph, Photographic-image, and Visual-illustration. Each frame was labeled independently by two annotators, resulting in an inter-annotator F1 score of 83.41%, indicating strong agreement. To ensure high-quality consensus annotations, a third expert reviewed and resolved all cases of disagreement through a conflict resolution process. To expand the dataset, a semi-supervised approach was employed to automatically annotate the remaining 3,000 frames, forming LVVO_3k. The complete dataset offers a valuable resource for developing and evaluating both supervised and semi-supervised methods for visual content detection in educational videos. The LVVO dataset is publicly available to support further research in this domain.

📄 PDF Abstract BibTeX arXiv:2506.13657

Code (2)

dipayan1109033/edu-video-visual-detection 공식 구현 pytorch
dipayan1109033/lvvo_dataset 공식 구현

Tasks

object-detectionObject Detection

Similar Papers 제목 키워드 기반

Unsupervised Audio-Visual Lecture Segmentation

2022-10-29 · Darshan Singh S, Anchit Gupta, C. V. Jawahar, Makarand Tapaswi

Over the last decade, online lecture videos have become increasingly popular and have experienced a meteoric rise during the pandemic. However, video-language research has primarily focused on instructional videos or mov…

NavigateOptical Character Recognition (OCR)Segmentation

Visual Summarization of Lecture Video Segments for Enhanced Navigation

2020-06-03 · Mohammad Rajiur Rahman, Jaspal Subhlok, Shishir Shah

Lecture videos are an increasingly important learning resource for higher education. However, the challenge of quickly finding the content of interest in a lecture video is an important limitation of this format. This pa…

Management

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

2025-04-20 · Enxin Song, Wenhao Chai, Weili Xu, Jianwen Xie 외

Recent advancements in language multimodal models (LMMs) for video have demonstrated their potential for understanding video content, yet the task of comprehending multi-discipline lectures remains largely unexplored. We…

MMLU

Generating Narrated Lecture Videos from Slides with Synchronized Highlights

2025-05-05 · Alexander Holmberg

Turning static slides into engaging video lectures takes considerable time and effort, requiring presenters to record explanations and visually guide their audience through the material. We introduce an end-to-end system…

Mathtext-to-speechText to Speech

PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos

2025-02-28 · Kangda Wei, Zhengyu Zhou, Bingqing Wang, Jun Araki 외

In recent years, online lecture videos have become an increasingly popular resource for acquiring new knowledge. Systems capable of effectively understanding/indexing lecture videos are thus highly desirable, enabling do…

Question AnsweringVideo Understanding