paper-with-me

홈 › Papers

Local Multi-Head Channel Self-Attention for Facial Expression Recognition

2021-11-14 · Roberto Pecoraro, Valerio Basile, Viviana Bono, Sara Gallo

Since the Transformer architecture was introduced in 2017 there has been many attempts to bring the self-attention paradigm in the field of computer vision. In this paper we propose a novel self-attention module that can be easily integrated in virtually every convolutional neural network and that is specifically designed for computer vision, the LHC: Local (multi) Head Channel (self-attention). LHC is based on two main ideas: first, we think that in computer vision the best way to leverage the self-attention paradigm is the channel-wise application instead of the more explored spatial attention and that convolution will not be replaced by attention modules like recurrent networks were in NLP; second, a local approach has the potential to better overcome the limitations of convolution than global attention. With LHC-Net we managed to achieve a new state of the art in the famous FER2013 dataset with a significantly lower complexity and impact on the "host" architecture in terms of computational cost when compared with the previous SOTA.

📄 PDF Abstract BibTeX arXiv:2111.07224

Code (1)

bodhis4ttva/lhc_net 공식 구현 tf

Tasks

Facial Expression RecognitionFacial Expression Recognition (FER)

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

Attention-based Neural Beamforming Layers for Multi-channel Speech Recognition

2021-05-12 · Bhargav Pulugundla, Yang Gao, Brian King, Gokce Keskin 외

Attention-based beamformers have recently been shown to be effective for multi-channel speech recognition. However, they are less capable at capturing local information. In this work, we propose a 2D Conv-Attention modul…

speech-recognitionSpeech Recognition

MILAAP: Mobile Link Allocation via Attention-based Prediction

2025-06-24 · Yung-Fu Chen, Anish Arora

Channel hopping (CS) communication systems must adapt to interference changes in the wireless network and to node mobility for maintaining throughput efficiency. Optimal scheduling requires up-to-date network state infor…

PredictionScheduling

DiNAT-IR: Exploring Dilated Neighborhood Attention for High-Quality Image Restoration

2025-07-23 · Hanzhou Liu, Binghan Li, Chengkai Liu, Mi Lu arxiv

Transformers, with their self-attention mechanisms for modeling long-range dependencies, have become a dominant paradigm in image restoration tasks. However, the high computational cost of self-attention limits scalabili…

Image Restoration

LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution

2024-09-05 · JeongSoo Kim, Jongho Nang, Junsuk Choe

Recent Vision Transformer (ViT)-based methods for Image Super-Resolution have demonstrated impressive performance. However, they suffer from significant complexity, resulting in high inference times and memory usage. Add…

GPUImage Super-ResolutionSuper-Resolution

ELSA: Enhanced Local Self-Attention for Vision Transformer

2021-12-23 · Jingkai Zhou, Pichao Wang, Fan Wang, Qiong Liu 외

Self-attention is powerful in modeling long-range dependencies, but it is weak in local finer-level feature learning. The performance of local self-attention (LSA) is just on par with convolution and inferior to dynamic …

Image ClassificationInstance SegmentationObject DetectionSemantic Segmentation