paper-with-me

홈 › Papers

The MSR-Video to Text Dataset with Clean Annotations

2021-02-12 · Haoran Chen, Jianmin Li, Simone Frintrop, Xiaolin Hu

Video captioning automatically generates short descriptions of the video content, usually in form of a single sentence. Many methods have been proposed for solving this task. A large dataset called MSR Video to Text (MSR-VTT) is often used as the benchmark dataset for testing the performance of the methods. However, we found that the human annotations, i.e., the descriptions of video contents in the dataset are quite noisy, e.g., there are many duplicate captions and many captions contain grammatical problems. These problems may pose difficulties to video captioning models for learning underlying patterns. We cleaned the MSR-VTT annotations by removing these problems, then tested several typical video captioning models on the cleaned dataset. Experimental results showed that data cleaning boosted the performances of the models measured by popular quantitative metrics. We recruited subjects to evaluate the results of a model trained on the original and cleaned datasets. The human behavior experiment demonstrated that trained on the cleaned dataset, the model generated captions that were more coherent and more relevant to the contents of the video clips.

📄 PDF Abstract BibTeX arXiv:2102.06448

Code (1)

WingsBrokenAngel/MSR-VTT-DataCleaning 공식 구현 tf

Tasks

SentenceVideo Captioning

Similar Papers 제목 키워드 기반

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

2024-11-13 · XiaoFeng Wang, Kang Zhao, Feng Liu, Jiayu Wang 외

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective…

Video Generation

Learning From Noisy Large-Scale Datasets With Minimal Supervision

2017-01-06 · CVPR 2017 7 · Andreas Veit, Neil Alldrin, Gal Chechik, Ivan Krasin 외

We present an approach to effectively use millions of images with noisy annotations in conjunction with a small subset of cleanly-annotated images to learn powerful image representations. One common approach to combine c…

DRIVE-C: A Controlled Corruption Dataset for Autonomous Driving

2026-05-10 · Shiva Aher arxiv

DRIVE-C is a controlled corruption dataset designed to evaluate visual perception robustness in autonomous driving systems. It is built from real-world forward-facing driving videos collected across daytime, nighttime, u…

Autonomous Driving

Collaborative Noisy Label Cleaner: Learning Scene-aware Trailers for Multi-modal Highlight Detection in Movies

2023-03-26 · CVPR 2023 1 · Bei Gan, Xiujun Shu, Ruizhi Qiao, Haoqian Wu 외

Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1) For different annotators, labeling hig…

Highlight DetectionLearning with noisy labelsScene Segmentation

Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation Networks

2019-10-01 · ICCV 2019 10 · Ziyi Liu, Le Wang, Qilin Zhang, Zhanning Gao 외

Weakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotation…

Action ClassificationAction LocalizationTemporal Action LocalizationWeakly Supervised Action Localization+1