paper-with-me

Papers

Video Summarization: Towards Entity-Aware Captions

2023-12-01 · Hammad A. Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Jialu Liu, Shih-Fu Chang

Existing popular video captioning benchmarks and models deal with generic captions devoid of specific person, place or organization named entities. In contrast, news videos present a challenging setting where the caption requires such named entities for meaningful summarization. As such, we propose the task of summarizing news video directly to entity-aware captions. We also release a large-scale dataset, VIEWS (VIdeo NEWS), to support research on this task. Further, we propose a method that augments visual information from videos with context retrieved from external world knowledge to generate entity-aware captions. We demonstrate the effectiveness of our approach on three video captioning models. We also show that our approach generalizes to existing news image captions dataset. With all the extensive experiments and insights, we believe we establish a solid basis for future research on this challenging task.

📄 PDF Abstract BibTeX arXiv:2312.02188

Code (1)

hayyubi/views 공식 구현

Tasks

Image CaptioningVideo CaptioningVideo SummarizationWorld Knowledge

Similar Papers 제목 키워드 기반

Identity-Aware Human-Object Interaction Motion Captioning

2026-08-21 · Yiming Wang, Yonghao Dang, Huilai Li, Jiawei Tu 외 arxiv

Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subje…

Motion Captioning

Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video

2024-05-14 · Tomoya Sugihara, Shuntaro Masuda, Ling Xiao, Toshihiko Yamasaki

Current video summarization methods rely heavily on supervised computer vision techniques, which demands time-consuming and subjective manual annotations. To overcome these limitations, we investigated self-supervised vi…

DiversitySupervised Video SummarizationVideo Summarization

Video Summarization with Large Language Models

2025-04-15 · CVPR 2025 1 · Min Jung Lee, Dayoung Gong, Minsu Cho

The exponential increase in video content poses significant challenges in terms of efficient navigation, search, and retrieval, thus requiring advanced video summarization techniques. Existing video summarization methods…

Large Language ModelVideo Summarization

Video Paragraph Captioning as a Text Summarization Task

2021-08-01 · ACL 2021 5 · Hui Liu, Xiaojun Wan

Video paragraph captioning aims to generate a set of coherent sentences to describe a video that contains several events. Most previous methods simplify this task by using ground-truth event segments. In this work, we pr…

SentenceText Summarization

ERA: Entity Relationship Aware Video Summarization with Wasserstein GAN

2021-09-06 · Guande Wu, Jianzhe Lin, Claudio T. Silva

Video summarization aims to simplify large scale video browsing by generating concise, short summaries that diver from but well represent the original video. Due to the scarcity of video annotations, recent progress for …

Unsupervised Video SummarizationVideo Summarization