Alignment-guided Temporal Attention for Video Action Recognition
Temporal modeling is crucial for various video learning tasks. Most recent approaches employ either factorized (2D+1D) or joint (3D) spatial-temporal operations to extract temporal contexts from the input frames. While the former is more efficient in computation, the latter often obtains better performance. In this paper, we attribute this to a dilemma between the sufficiency and the efficiency of interactions among various positions in different frames. These interactions affect the extraction of task-relevant information shared among frames. To resolve this issue, we prove that frame-by-frame alignments have the potential to increase the mutual information between frame representations, thereby including more task-relevant information to boost effectiveness. Then we propose Alignment-guided Temporal Attention (ATA) to extend 1-dimensional temporal attention with parameter-free patch-level alignments between neighboring frames. It can act as a general plug-in for image backbones to conduct the action recognition task without any model-specific design. Extensive experiments on multiple benchmarks demonstrate the superiority and generality of our module.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAttributeTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Video-to-Task Learning via Motion-Guided Attention for Few-Shot Action Recognition
In recent years, few-shot action recognition has achieved remarkable performance through spatio-temporal relation modeling. Although a wide range of spatial and temporal alignment modules have been proposed, they primari…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionMEMO: Memory-Guided Diffusion for Expressive Talking Video Generation
Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency…
Portrait AnimationVideo GenerationMF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment
The rapid proliferation of online video content necessitates effective video summarization techniques. Traditional methods, often relying on a single modality (typically visual), struggle to capture the full semantic ric…
Video SummarizationSwap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention. Generating videos guided by text instructions poses significant challenges, such as modeling the co…
Image GenerationText to Image GenerationText-to-Image GenerationText-to-Video Generation+2Beyond Alignment: Blind Video Face Restoration via Parsing-Guided Temporal-Coherent Transformer
Multiple complex degradations are coupled in low-quality video faces in the real world. Therefore, blind video face restoration is a highly challenging ill-posed problem, requiring not only hallucinating high-fidelity de…
Face ParsingSemantic ParsingVideo Temporal Consistency