Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation
We propose \textbf{Compressed Video Aggregator} (CVA), a lightweight micro-video recommendation module that decouples video information from preference learning. CVA first summarizes frozen VFM frame embeddings into a semantic-consensus anchor through masked mean pooling, projects this anchor into a compact latent space, and refines the projected representation with residual self-attention and feedforward blocks before producing a single video embedding for the recommender. Due to the redundancy in the frame count of the original benchmark and its overly coarse sampling, we used titles to re-select key frames based on CLIP. Experiments on MicroLens and Short-Video show consistent gains with orders-of-magnitude reductions in training time and GPU memory, and re-selected frames can further enhance the performance of all methods, including CVA. Furthermore, we also discussed the impact of several scenarios involving erroneous titles on our method.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
COMISR: Compression-Informed Video Super-Resolution
Most video super-resolution methods focus on restoring high-resolution video frames from low-resolution videos without taking into account compression. However, most videos on the web or mobile devices are compressed, an…
Super-ResolutionVideo Super-ResolutionDeep Learning based Full-reference and No-reference Quality Assessment Models for Compressed UGC Videos
In this paper, we propose a deep learning based video quality assessment (VQA) framework to evaluate the quality of the compressed user's generated content (UGC) videos. The proposed VQA framework consists of three modul…
regressionVideo Quality AssessmentMulti-Attention Network for Compressed Video Referring Object Segmentation
Referring video object segmentation aims to segment the object referred by a given language expression. Existing works typically require compressed video bitstream to be decoded to RGB frames before being segmented, whic…
ObjectReferring Expression SegmentationReferring Video Object SegmentationSegmentation+3CCLAP: Controllable Chinese Landscape Painting Generation via Latent Diffusion Model
With the development of deep generative models, recent years have seen great success of Chinese landscape painting generation. However, few works focus on controllable Chinese landscape painting generation due to the lac…
Chinese Landscape Painting GenerationTask Oriented Video Coding: A Survey
Video coding technology has been continuously improved for higher compression ratio with higher resolution. However, the state-of-the-art video coding standards, such as H.265/HEVC and Versatile Video Coding, are still d…
Survey