paper-with-me

홈 › Papers

Less Is More: Picking Informative Frames for Video Captioning

2018-03-05 · ECCV 2018 9 · Yangyu Chen, Shuhui Wang, Weigang Zhang, Qingming Huang

In video captioning task, the best practice has been achieved by attention-based models which associate salient visual components with sentences in the video. However, existing study follows a common procedure which includes a frame-level appearance modeling and motion modeling on equal interval frame sampling, which may bring about redundant visual information, sensitivity to content noise and unnecessary computation cost. We propose a plug-and-play PickNet to perform informative frame picking in video captioning. Based on a standard Encoder-Decoder framework, we develop a reinforcement-learning-based procedure to train the network sequentially, where the reward of each frame picking action is designed by maximizing visual diversity and minimizing textual discrepancy. If the candidate is rewarded, it will be selected and the corresponding latent representation of Encoder-Decoder will be updated for future trials. This procedure goes on until the end of the video sequence. Consequently, a compact frame subset can be selected to represent the visual information and perform video captioning without performance degradation. Experiment results shows that our model can use 6-8 frames to achieve competitive performance across popular benchmarks.

📄 PDF Abstract BibTeX arXiv:1803.01457

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDiversityReinforcement LearningVideo Captioning

Similar Papers 제목 키워드 기반

OCSampler: Compressing Videos to One Clip with Single-step Sampling

2022-01-12 · CVPR 2022 1 · Jintao Lin, Haodong Duan, Kai Chen, Dahua Lin 외

In this paper, we propose a framework named OCSampler to explore a compact yet effective video representation with one short clip for efficient video recognition. Recent works prefer to formulate frame sampling as a sequ…

GPUVideo Recognition

PEEK: Picking Essential frames via Efficient Knowledge distillation

2026-05-29 · Killian Steunou, Anas Filali Razzouki, Khalil Guetari, Mounîm A. El-Yacoubi 외 arxiv

Video-language models can process only a limited number of frames, making frame selection a key bottleneck for efficient video captioning. Most captioning pipelines still rely on uniform sampling, which is computationall…

Knowledge DistillationVideo Captioning

Human activity recognition using improved dynamic image

2020-11-15 · IET Image Processing 2020 11 · Mohammadreza Riahi, Mohammad Eslami, Seyed Hamid Safavi, Farah Torkamani Azar

In action recognition, the dynamic image (DI) approach is recently proposed to code a video signal to a still image. Since DI descriptor is strongly dependent on first frames, it cannot extract dynamics that do not occur…

Action RecognitionActivity RecognitionHuman Activity Recognition

Active Learning for Video Classification with Frame Level Queries

2023-07-10 · International Joint Conference on Neural Networks (IJCNN) 2023 8 · Debanjan Goswami, Shayok Chakraborty

Deep learning algorithms have pushed the boundaries of computer vision research and have depicted commendable performance in a variety of applications. However, training a robust deep neural network necessitates a large …

Active LearningClassificationVideo Classification

Autofluorescence Bronchoscopy Video Analysis for Lesion Frame Detection

2023-03-21 · Qi Chang, Rebecca Bascom, Jennifer Toth, Danish Ahmad 외

Because of the significance of bronchial lesions as indicators of early lung cancer and squamous cell carcinoma, a critical need exists for early detection of bronchial lesions. Autofluorescence bronchoscopy (AFB) is a p…

Lesion Detection