paper-with-me

홈 › Papers

What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations

2025-02-12 · Dongqi Liu, Chenxi Whitehouse, Xi Yu, Louis Mahon, Rohit Saxena, Zheng Zhao, Yifu Qiu, Mirella Lapata, Vera Demberg

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains. VISTA contains 18,599 recorded AI conference presentations paired with their corresponding paper abstracts. We benchmark the performance of state-of-the-art large models and apply a plan-based framework to better capture the structured nature of abstracts. Both human and automated evaluations confirm that explicit planning enhances summary quality and factual consistency. However, a considerable gap remains between models and human performance, highlighting the challenges of scientific video summarization.

📄 PDF Abstract BibTeX arXiv:2502.08279

Code (1)

dongqi-me/vista 공식 구현 pytorch

Tasks

Text SummarizationVideo Summarization

Similar Papers 제목 키워드 기반

TalkSumm: A Dataset and Scalable Annotation Method for Scientific Paper Summarization Based on Conference Talks

2019-06-04 · ACL 2019 7 · Guy Lev, Michal Shmueli-Scheuer, Jonathan Herzig, Achiya Jerbi 외

Currently, no large-scale training data is available for the task of scientific paper summarization. In this paper, we propose a novel method that automatically generates summaries for scientific papers, by utilizing vid…

Dial2Desc: End-to-end Dialogue Description Generation

2018-11-01 · Haojie Pan, Junpei Zhou, Zhou Zhao, Yan Liu 외

We first propose a new task named Dialogue Description (Dial2Desc). Unlike other existing dialogue summarization tasks such as meeting summarization, we do not maintain the natural flow of a conversation but describe an …

DescriptiveMeeting Summarization

Generation of Multimedia Artifacts: An Extractive Summarization-based Approach

2015-08-13 · Paulo Figueiredo, Marta Aparício, David Martins de Matos, Ricardo Ribeiro

We explore methods for content selection and address the issue of coherence in the context of the generation of multimedia artifacts. We use audio and video to present two case studies: generation of film tributes, and l…

DiversityExtractive Summarization

What are they talking about? Benchmarking Large Language Models for Knowledge-Grounded Discussion Summarization

2025-05-18 · Weixiao Zhou, Junnan Zhu, Gengyao Li, Xianfu Cheng 외

In this work, we investigate the performance of LLMs on a new task that requires combining discussion with background knowledge for summarization. This aims to address the limitation of outside observer confusion in exis…

Benchmarking

Stealing Creator's Workflow: A Creator-Inspired Agentic Framework with Iterative Feedback Loop for Improved Scientific Short-form Generation

2025-04-26 · Jong Inn Park, Maanas Taneja, Qianwen Wang, Dongyeop Kang

Generating engaging, accurate short-form videos from scientific papers is challenging due to content complexity and the gap between expert authors and readers. Existing end-to-end methods often suffer from factual inaccu…

FormVideo Generation