paper-with-me

홈 › Papers

Script-to-Slide Grounding: Grounding Script Sentences to Slide Objects for Automatic Instructional Video Generation

2026-03-14 · Rena Suzuki, Masato Kikuchi, Tadachika Ozono arxiv

While slide-based videos augmented with visual effects are widely utilized in education and research presentations, the video editing process -- particularly applying visual effects to ground spoken content to slide objects -- remains highly labor-intensive. This study aims to develop a system that automatically generates such instructional videos from slides and corresponding scripts. As a foundational step, this paper proposes and formulates Script-to-Slide Grounding (S2SG), defined as the task of grounding script sentences to their corresponding slide objects. Furthermore, as an initial step, we propose ``Text-S2SG,'' a method that utilizes a large language model (LLM) to perform this grounding task for text objects. Our experiments demonstrate that the proposed method achieves high performance (F1-score: 0.924). The contribution of this work is the formalization of a previously implicit slide-based video editing process into a computable task, thereby paving the way for its automation.

📄 PDF Abstract BibTeX arXiv:2603.16931

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Generating Descriptions with Grounded and Co-Referenced People

2017-04-05 · CVPR 2017 7 · Anna Rohrbach, Marcus Rohrbach, Siyu Tang, Seong Joon Oh 외

Learning how to generate descriptions of images or videos received major interest both in the Computer Vision and Natural Language Processing communities. While a few works have proposed to learn a grounding during the g…

Multi-Sentence Grounding for Long-term Instructional Video

2023-12-21 · Zeqian Li, Qirui Chen, Tengda Han, Ya zhang 외

In this paper, we aim to establish an automatic, scalable pipeline for denoising the large-scale instructional dataset and construct a high-quality video-text dataset with multiple descriptive steps supervision, named Ho…

DenoisingDescriptiveLanguage ModellingLarge Language Model+3

Grounding Action Descriptions in Videos

2013-01-01 · TACL 2013 1 · Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater 외

Recent work has shown that the integration of visual information into text-based models can substantially improve model predictions, but so far only visual information extracted from static images has been used. In this …

Semantic Textual SimilarityVideo Understanding

MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions

2021-12-01 · CVPR 2022 1 · Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron 외

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at asse…

Moment RetrievalNatural Language Moment Retrieval

DeepSlide: From Artifacts to Presentation Delivery

2026-04-01 · Ming Yang, Zhiwei Zhang, Jiahang Li, Haoseng Liu 외 arxiv

Presentations are a primary medium for scholarly communication, yet most AI slide generators optimize the artifact (a visually plausible deck) while under-optimizing the delivery process (pacing, narrative, and presentat…