paper-with-me

Papers

Fine-grained Video Categorization with Redundancy Reduction Attention

2018-10-26 · ECCV 2018 9 · Chen Zhu, Xiao Tan, Feng Zhou, Xiao Liu, Kaiyu Yue, Errui Ding, Yi Ma

For fine-grained categorization tasks, videos could serve as a better source than static images as videos have a higher chance of containing discriminative patterns. Nevertheless, a video sequence could also contain a lot of redundant and irrelevant frames. How to locate critical information of interest is a challenging task. In this paper, we propose a new network structure, known as Redundancy Reduction Attention (RRA), which learns to focus on multiple discriminative patterns by sup- pressing redundant feature channels. Specifically, it firstly summarizes the video by weight-summing all feature vectors in the feature maps of selected frames with a spatio-temporal soft attention, and then predicts which channels to suppress or to enhance according to this summary with a learned non-linear transform. Suppression is achieved by modulating the feature maps and threshing out weak activations. The updated feature maps are then used in the next iteration. Finally, the video is classified based on multiple summaries. The proposed method achieves out- standing performances in multiple video classification datasets. Further- more, we have collected two large-scale video datasets, YouTube-Birds and YouTube-Cars, for future researches on fine-grained video categorization. The datasets are available at http://www.cs.umd.edu/~chenzhu/fgvc.

📄 PDF Abstract BibTeX arXiv:1810.11189

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionVideo Classification

Similar Papers 제목 키워드 기반

Exploring Fine-Grained Audiovisual Categorization with the SSW60 Dataset

2022-07-21 · Grant van Horn, Rui Qian, Kimberly Wilber, Hartwig Adam 외

We present a new benchmark dataset, Sapsucker Woods 60 (SSW60), for advancing research on audiovisual fine-grained categorization. While our community has made great strides in fine-grained visual categorization on image…

Fine-Grained Visual CategorizationVideo Classification

Towards Long Video Understanding via Fine-detailed Video Story Generation

2024-12-09 · Zeng You, Zhiquan Wen, Yaofo Chen, Xin Li 외

Long video understanding has become a critical task in computer vision, driving advancements across numerous applications from surveillance to content retrieval. Existing video understanding methods suffer from two chall…

Story GenerationVideo Understanding

R2-Trans:Fine-Grained Visual Categorization with Redundancy Reduction

2022-04-21 · Yu Wang, Shuo Ye, Shujian Yu, Xinge You

Fine-grained visual categorization (FGVC) aims to discriminate similar subcategories, whose main challenge is the large intraclass diversities and subtle inter-class differences. Existing FGVC methods usually select disc…

Fine-Grained Visual Categorization

Similarity Comparisons for Interactive Fine-Grained Categorization

2014-06-01 · CVPR 2014 6 · Catherine Wah, Grant van Horn, Steve Branson, Subhransu Maji 외

Current human-in-the-loop fine-grained visual categorization systems depend on a predefined vocabulary of attributes and parts, usually determined by experts. In this work, we move away from that expert-driven and attrib…

AttributeFine-Grained Visual CategorizationGeneral ClassificationImage Retrieval+1

OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs

2025-11-15 · Feng Chen, Yefei He, Shaoxuan He, Yuanyu He 외 arxiv

Existing sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge the training-inference gap and lack the c…

Semantic Similarity