Guiding the Flowing of Semantics: Interpretable Video Captioning via POS Tag
In the current video captioning models, the video frames are collected in one network and the semantics are mixed into one feature, which not only increase the difficulty of the caption decoding, but also decrease the interpretability of the captioning models. To address these problems, we propose an Adaptive Semantic Guidance Network (ASGN), which instantiates the whole video semantics to different POS-aware semantics with the supervision of part of speech (POS) tag. In the encoding process, the POS tag activates the related neurons and parses the whole semantic information into corresponding encoded video representations. Furthermore, the potential of the model is stimulated by the POS-aware video features. In the decoding process, the related video features of noun and verb are used as the supervision to construct a new adaptive attention model which can decide whether to attend to the video feature or not. With the explicit improving of the interpretability of the network, the learning process is more transparent and the results are more predictable. Extensive experiments demonstrate the effectiveness of our model when compared with state-of-the-art models.
Code (0)
등록된 구현이 없습니다.
Tasks
POSTAGVideo CaptioningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
IcoCap: Improving Video Captioning by Compounding Images
Video captioning is a more challenging task compared to image captioning, primarily due to differences in content density. Video data contains redundant visual content, making it difficult for captioners to generalize di…
Image CaptioningVideo CaptioningParallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation. In this domain, autoregressive video large language models have em…
Dense Video CaptioningAn Attempt towards Interpretable Audio-Visual Video Captioning
Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory and visual contents are equally important…
Audio captioningAudio-Visual Video CaptioningImage CaptioningSentence+2Syntax Customized Video Captioning by Imitating Exemplar Sentences
Enhancing the diversity of sentences to describe video contents is an important problem arising in recent video captioning research. In this paper, we explore this problem from a novel perspective of customizing video ca…
DecoderDiversitySentencevalid+1Guided Attention for Interpretable Motion Captioning
Diverse and extensive work has recently been conducted on text-conditioned human motion generation. However, progress in the reverse direction, motion captioning, has seen less comparable advancement. In this paper, we i…
Action LocalizationMotion CaptioningMotion GenerationSpatio-Temporal Video Grounding+1