paper-with-me

Papers

Interweaving Insights: High-Order Feature Interaction for Fine-Grained Visual Recognition

2024-10-20 · International Journal of Computer Vision 2024 10 · Arindam Sikdar, Yonghuai Liu, Siddhardha Kedarisetty, Yitian Zhao, Amr Ahmed & Ardhendu Behera

This paper presents a novel approach for Fine-Grained Visual Classification (FGVC) by exploring Graph Neural Networks (GNNs) to facilitate high-order feature interactions, with a specific focus on constructing both inter- and intra-region graphs. Unlike previous FGVC techniques that often isolate global and local features, our method combines both features seamlessly during learning via graphs. Inter-region graphs capture long-range dependencies to recognize global patterns, while intra-region graphs delve into finer details within specific regions of an object by exploring high-dimensional convolutional features. A key innovation is the use of shared GNNs with an attention mechanism coupled with the Approximate Personalized Propagation of Neural Predictions (APPNP) message-passing algorithm, enhancing information propagation efficiency for better discriminability and simplifying the model architecture for computational efficiency. Additionally, the introduction of residual connections improves performance and training stability. Comprehensive experiments showcase state-of-the-art results on benchmark FGVC datasets, affirming the efficacy of our approach. This work underscores the potential of GNN in modeling high-level feature interactions, distinguishing it from previous FGVC methods that typically focus on singular aspects of feature representation. Our source code is available at https://github.com/Arindam-1991/I2-HOFI.

📄 PDF Abstract BibTeX

Code (1)

arindam-1991/i2-hofi 공식 구현 tf

Tasks

Computational EfficiencyFine-Grained Image ClassificationFine-Grained Visual Recognition

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Multistep feature aggregation framework for salient object detection

2022-11-12 · Xiaogang Liu Shuang Song

Recent works on salient object detection have made use of multi-scale features in a way such that high-level features and low-level features can collaborate in locating salient objects. Many of the previous methods have …

Objectobject-detectionObject DetectionSalient Object Detection

Play with Emotion: Affect-Driven Reinforcement Learning

2022-08-26 · Matthew Barthet, Ahmed Khalifa, Antonios Liapis, Georgios N. Yannakakis

This paper introduces a paradigm shift by viewing the task of affect modeling as a reinforcement learning (RL) process. According to the proposed paradigm, RL agents learn a policy (i.e. affective interaction) by attempt…

Decision Makingreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Interpretable Click-Through Rate Prediction through Hierarchical Attention

2020-01-20 · International Conference on Web Search and Data Mining 2020 1 · Zeyu Li, Wei Cheng, Yang Chen, Haifeng Chen 외

Click-through rate (CTR) prediction is a critical task in online advertising and marketing. For this problem, existing approaches, with shallow or deep architectures, have three major drawbacks. First, they typically lac…

Click-Through Rate PredictionMarketingPrediction

"Previously on ..." From Recaps to Story Summarization

2024-05-19 · Aditya Kumar Singh, Dhruv Srivastava, Makarand Tapaswi

We introduce multimodal story summarization by leveraging TV episode recaps - short video sequences interweaving key story moments from previous episodes to bring viewers up to speed. We propose PlotSnap, a dataset featu…

Video Summarization

Previously on ... From Recaps to Story Summarization

2024-01-01 · CVPR 2024 1 · Aditya Kumar Singh, Dhruv Srivastava, Makarand Tapaswi

We introduce multimodal story summarization by leveraging TV episode recaps - short video sequences interweaving key story moments from previous episodes to bring viewers up to speed. We propose PlotSnap a dataset fe…

Video Summarization