paper-with-me

홈 › Papers

A Semantics-Assisted Video Captioning Model Trained with Scheduled Sampling

2019-08-31 · Haoran Chen, Ke Lin, Alexander Maye, Jianming Li, Xiaolin Hu

Given the features of a video, recurrent neural networks can be used to automatically generate a caption for the video. Existing methods for video captioning have at least three limitations. First, semantic information has been widely applied to boost the performance of video captioning models, but existing networks often fail to provide meaningful semantic features. Second, the Teacher Forcing algorithm is often utilized to optimize video captioning models, but during training and inference, different strategies are applied to guide word generation, leading to poor performance. Third, current video captioning models are prone to generate relatively short captions that express video contents inappropriately. Toward resolving these three problems, we suggest three corresponding improvements. First of all, we propose a metric to compare the quality of semantic features, and utilize appropriate features as input for a semantic detection network (SDN) with adequate complexity in order to generate meaningful semantic features for videos. Then, we apply a scheduled sampling strategy that gradually transfers the training phase from a teacher-guided manner toward a more self-teaching manner. Finally, the ordinary logarithm probability loss function is leveraged by sentence length so that the inclination of generating short sentences is alleviated. Our model achieves better results than previous models on the YouTube2Text dataset and is competitive with the previous best model on the MSR-VTT dataset.

📄 PDF Abstract BibTeX arXiv:1909.00121

Code (2)

WingsBrokenAngel/Semantics-AssistedVideoCaptioning 공식 구현 tf
WingsBrokenAngel/Semantics-AssistedVideoCaptioningModelTrainedwithScheduledSamplingStrategy tf

Tasks

SentenceVideo Captioning

Similar Papers 제목 키워드 기반

PIC 4th Challenge: Semantic-Assisted Multi-Feature Encoding and Multi-Head Decoding for Dense Video Captioning

2022-07-06 · Yifan Lu, Ziqi Zhang, Yuxin Chen, Chunfeng Yuan 외

The task of Dense Video Captioning (DVC) aims to generate captions with timestamps for multiple events in one video. Semantic information plays an important role for both localization and description of DVC. We present a…

Dense Video CaptioningVideo Captioning

IcoCap: Improving Video Captioning by Compounding Images

2023-10-05 · IEEE Transactions on Multimedia 2023 10 · Yuanzhi Liang, Linchao Zhu, Xiaohan Wang, Yi Yang

Video captioning is a more challenging task compared to image captioning, primarily due to differences in content density. Video data contains redundant visual content, making it difficult for captioners to generalize di…

Image CaptioningVideo Captioning

Guiding the Flowing of Semantics: Interpretable Video Captioning via POS Tag

2019-11-01 · IJCNLP 2019 11 · Xinyu Xiao, Lingfeng Wang, Bin Fan, Shinming Xiang 외

In the current video captioning models, the video frames are collected in one network and the semantics are mixed into one feature, which not only increase the difficulty of the caption decoding, but also decrease the in…

POSTAGVideo Captioning

Syntax Customized Video Captioning by Imitating Exemplar Sentences

2021-12-02 · Yitian Yuan, Lin Ma, Wenwu Zhu

Enhancing the diversity of sentences to describe video contents is an important problem arising in recent video captioning research. In this paper, we explore this problem from a novel perspective of customizing video ca…

DecoderDiversitySentencevalid+1

Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction

2025-11-13 · Mingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li 외 arxiv

Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competiti…

Dense Video CaptioningCross-Modal Retrieval