paper-with-me

홈 › Papers

Watch It Twice: Video Captioning with a Refocused Video Encoder

2019-07-21 · Xiangxi Shi, Jianfei Cai, Shafiq Joty, Jiuxiang Gu

With the rapid growth of video data and the increasing demands of various applications such as intelligent video search and assistance toward visually-impaired people, video captioning task has received a lot of attention recently in computer vision and natural language processing fields. The state-of-the-art video captioning methods focus more on encoding the temporal information, while lack of effective ways to remove irrelevant temporal information and also neglecting the spatial details. However, the current RNN encoding module in single time order can be influenced by the irrelevant temporal information, especially the irrelevant temporal information is at the beginning of the encoding. In addition, neglecting spatial information will lead to the relationship confusion of the words and detailed loss. Therefore, in this paper, we propose a novel recurrent video encoding method and a novel visual spatial feature for the video captioning task. The recurrent encoding module encodes the video twice with the predicted key frame to avoid the irrelevant temporal information often occurring at the beginning and the end of a video. The novel spatial features represent the spatial information in different regions of a video and enrich the details of a caption. Experiments on two benchmark datasets show superior performance of the proposed method.

📄 PDF Abstract BibTeX arXiv:1907.12905

Code (0)

등록된 구현이 없습니다.

Tasks

Video Captioning

Similar Papers 제목 키워드 기반

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization

2025-07-02 · Jiyang Tang, Hengyi Li, Yifan Du, Wayne Xin Zhao arxiv

Although video multimodal large language models (video MLLMs) have achieved substantial progress in video captioning tasks, it remains challenging to adjust the focal emphasis of video captions according to human prefere…

Video Captioning

Small Clips, Big Gains: Learning Long-Range Refocused Temporal Information for Video Super-Resolution

2025-05-04 · Xingyu Zhou, Wei Long, Jingbo Lu, Shiyin Jiang 외

Video super-resolution (VSR) can achieve better performance compared to single image super-resolution by additionally leveraging temporal information. In particular, the recurrent-based VSR model exploits long-range temp…

Computational EfficiencyImage Super-ResolutionSuper-ResolutionVideo Super-Resolution

Video Captioning with Boundary-aware Hierarchical Language Decoding and Joint Video Prediction

2018-07-08 · Xiangxi Shi, Jianfei Cai, Jiuxiang Gu, Shafiq Joty

The explosion of video data on the internet requires effective and efficient technology to generate captions automatically for people who are not able to watch the videos. Despite the great progress of video captioning r…

DecoderLanguage ModelingLanguage ModellingSentence+3

Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning

2018-04-15 · NAACL 2018 6 · Xin Wang, Yuan-Fang Wang, William Yang Wang

A major challenge for video captioning is to combine audio and visual cues. Existing multi-modal fusion methods have shown encouraging results in video understanding. However, the temporal structures of multiple modaliti…

Video CaptioningVideo Understanding

SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries

2023-04-10 · Hassan Mkhallati, Anthony Cioppa, Silvio Giancola, Bernard Ghanem 외

Soccer is more than just a game - it is a passion that transcends borders and unites people worldwide. From the roar of the crowds to the excitement of the commentators, every moment of a soccer match is a thrill. Yet, w…

Dense Video CaptioningVideo Captioning