paper-with-me

홈 › Papers

A Human-Annotated Video Dataset for Training and Evaluation of 360-Degree Video Summarization Methods

2024-06-05 · Ioannis Kontostathis, Evlampios Apostolidis, Vasileios Mezaris

In this paper we introduce a new dataset for 360-degree video summarization: the transformation of 360-degree video content to concise 2D-video summaries that can be consumed via traditional devices, such as TV sets and smartphones. The dataset includes ground-truth human-generated summaries, that can be used for training and objectively evaluating 360-degree video summarization methods. Using this dataset, we train and assess two state-of-the-art summarization methods that were originally proposed for 2D-video summarization, to serve as a baseline for future comparisons with summarization methods that are specifically tailored to 360-degree video. Finally, we present an interactive tool that was developed to facilitate the data annotation process and can assist other annotation activities that rely on video fragment selection.

📄 PDF Abstract BibTeX arXiv:2406.02991

Code (1)

idt-iti/360-vsumm 공식 구현 pytorch

Tasks

Video Summarization

Similar Papers 제목 키워드 기반

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

2025-10-09 · Nikhil Reddy Varimalla, Yunfei Xu, Arkadiy Saakyan, Meng Fan Wang 외 arxiv

As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts. To advance cultural norm awareness evaluation in VideoLLMs, we introduce Video…

HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models

2025-02-28 · Xiao Wang, Jingyun Hua, WeiHong Lin, Yuanxing Zhang 외

Recent Multi-modal Large Language Models (MLLMs) have made great progress in video understanding. However, their performance on videos involving human actions is still limited by the lack of high-quality data. To address…

Action UnderstandingText-to-Video GenerationVideo GenerationVideo Understanding

The AVA-Kinetics Localized Human Actions Video Dataset

2020-05-01 · Ang Li, Meghana Thotakuri, David A. Ross, João Carreira 외

This paper describes the AVA-Kinetics localized human actions video dataset. The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation protocol, and extending the original AVA d…

Action Classification

Tencent-MVSE: A Large-Scale Benchmark Dataset for Multi-Modal Video Similarity Evaluation

2022-01-01 · CVPR 2022 1 · Zhaoyang Zeng, Yongsheng Luo, Zhenhua Liu, Fengyun Rao 외

Multi-modal video similarity evaluation is important for video recommendation systems such as video de-duplication, relevance matching, ranking, and diversity control. However, there still lacks a benchmark dataset t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityRecommendation Systems+3

DeVAn: Dense Video Annotation for Video-Language Models

2023-10-08 · Tingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan 외

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset…

RetrievalSentenceVideo Summarization