paper-with-me

Papers

Not All Words are Equal: Video-specific Information Loss for Video Captioning

2019-01-01 · Jiarong Dong, Ke Gao, Xiaokai Chen, Junbo Guo, Juan Cao, Yongdong Zhang

An ideal description for a given video should fix its gaze on salient and representative content, which is capable of distinguishing this video from others. However, the distribution of different words is unbalanced in video captioning datasets, where distinctive words for describing video-specific salient objects are far less than common words such as 'a' 'the' and 'person'. The dataset bias often results in recognition error or detail deficiency of salient but unusual objects. To address this issue, we propose a novel learning strategy called Information Loss, which focuses on the relationship between the video-specific visual content and corresponding representative words. Moreover, a framework with hierarchical visual representations and an optimized hierarchical attention mechanism is established to capture the most salient spatial-temporal visual information, which fully exploits the potential strength of the proposed learning strategy. Extensive experiments demonstrate that the ingenious guidance strategy together with the optimized architecture outperforms state-of-the-art video captioning methods on MSVD with CIDEr score 87.5, and achieves superior CIDEr score 47.7 on MSR-VTT. We also show that our Information Loss is generic which improves various models by significant margins.

📄 PDF Abstract BibTeX arXiv:1901.00097

Code (0)

등록된 구현이 없습니다.

Tasks

AllVideo Captioning

Similar Papers 제목 키워드 기반

Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

2023-11-15 · WonJun Moon, Sangeek Hyun, SuBeen Lee, Jae-Pil Heo

Temporal Grounding is to identify specific moments or highlights from a video corresponding to textual descriptions. Typical approaches in temporal grounding treat all video clips equally during the encoding process rega…

Highlight DetectionMoment RetrievalNatural Language Moment RetrievalRepresentation Learning+1

Hierarchical Local-Global Transformer for Temporal Sentence Grounding

2022-08-31 · Xiang Fang, Daizong Liu, Pan Zhou, Zichuan Xu 외

This paper studies the multimedia problem of temporal sentence grounding (TSG), which aims to accurately determine the specific video segment in an untrimmed video according to a given sentence query. Traditional TSG met…

SentenceTemporal Sentence Grounding

Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses

2019-06-01 · CVPR 2019 6 · Jing Shi, Jia Xu, Boqing Gong, Chenliang Xu

We invest the problem of weakly-supervised video grounding, where only video-level sentences are provided. This is a challenging task, and previous Multi-Instance Learning (MIL) based image grounding methods turn to fai…

AllClusteringSentenceVideo Grounding

Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning

2017-06-05 · Jingkuan Song, Zhao Guo, Lianli Gao, Wu Liu 외

Recent progress has been made in using attention based encoder-decoder framework for video captioning. However, most existing decoders apply the attention mechanism to every generated word including both visual words (e.…

Caption GenerationDecoderLanguage ModelingLanguage Modelling+1

A Video-Aware FEC-Based Unequal Loss Protection System for Video Streaming over RTP

2024-02-07 · César Díaz, Julián Cabrera, Fernando Jaureguizar, Narciso García

A video-aware unequal loss protection (ULP) system for protecting RTP video streaming in bursty packet loss networks is proposed. Considering the relevance of the frame, the state of the channel, and the bitrate constrai…