paper-with-me

홈 › Papers

SnapCap: Efficient Snapshot Compressive Video Captioning

2024-01-10 · JianQiao Sun, Yudi Su, Hao Zhang, Ziheng Cheng, Zequn Zeng, Zhengjue Wang, Bo Chen, Xin Yuan

Video Captioning (VC) is a challenging multi-modal task since it requires describing the scene in language by understanding various and complex videos. For machines, the traditional VC follows the "imaging-compression-decoding-and-then-captioning" pipeline, where compression is pivot for storage and transmission. However, in such a pipeline, some potential shortcomings are inevitable, i.e., information redundancy resulting in low efficiency and information loss during the sampling process for captioning. To address these problems, in this paper, we propose a novel VC pipeline to generate captions directly from the compressed measurement, which can be captured by a snapshot compressive sensing camera and we dub our model SnapCap. To be more specific, benefiting from the signal simulation, we have access to obtain abundant measurement-video-annotation data pairs for our model. Besides, to better extract language-related visual representations from the compressed measurement, we propose to distill the knowledge from videos via a pre-trained CLIP with plentiful language-vision associations to guide the learning of our SnapCap. To demonstrate the effectiveness of SnapCap, we conduct experiments on two widely-used VC datasets. Both the qualitative and quantitative results verify the superiority of our pipeline over conventional VC pipelines. In particular, compared to the "caption-after-reconstruction" methods, our SnapCap can run at least 3$\times$ faster, and achieve better caption results.

📄 PDF Abstract BibTeX arXiv:2401.04903

Code (0)

등록된 구현이 없습니다.

Tasks

Compressive SensingVideo Captioning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Reinforcement Learning for Adaptive Video Compressive Sensing

2021-05-18 · Sidi Lu, Xin Yuan, Aggelos K Katsaggelos, Weisong Shi

We apply reinforcement learning to video compressive sensing to adapt the compression ratio. Specifically, video snapshot compressive imaging (SCI), which captures high-speed video using a low-speed camera is considered …

Compressive Sensingobject-detectionObject Detectionreinforcement-learning+3

Event-Enhanced Snapshot Compressive Videography at 10K FPS

2024-04-11 · Bo Zhang, Jinli Suo, Qionghai Dai

Video snapshot compressive imaging (SCI) encodes the target dynamic scene compactly into a snapshot and reconstructs its high-speed frame sequence afterward, greatly reducing the required data footprint and transmission …

Video Frame Interpolation

Key frames assisted hybrid encoding for photorealistic compressive video sensing

2022-07-26 · Honghao Huang, Jiajie Teng, Yu Liang, Chengyang Hu 외

Snapshot compressive imaging (SCI) encodes high-speed scene video into a snapshot measurement and then computationally makes reconstructions, allowing for efficient high-dimensional data acquisition. Numerous algorithms,…

Optical Flow Estimation

Rank Minimization for Snapshot Compressive Imaging

2018-07-20 · Yang Liu, Xin Yuan, Jinli Suo, David J. Brady 외

Snapshot compressive imaging (SCI) refers to compressive imaging systems where multiple frames are mapped into a single measurement, with video compressive imaging and hyperspectral compressive imaging as two representat…

Snapshot Compressive Imaging: Principle, Implementation, Theory, Algorithms and Applications

2021-03-07 · Xin Yuan, David J. Brady, Aggelos K. Katsaggelos

Capturing high-dimensional (HD) data is a long-term challenge in signal processing and related fields. Snapshot compressive imaging (SCI) uses a two-dimensional (2D) detector to capture HD ($\ge3$D) data in a {\em snapsh…