paper-with-me

Papers

MAFNet: A Multi-Attention Fusion Network for RGB-T Crowd Counting

2022-08-14 · PengYu Chen, Junyu Gao, Yuan Yuan, Qi Wang

RGB-Thermal (RGB-T) crowd counting is a challenging task, which uses thermal images as complementary information to RGB images to deal with the decreased performance of unimodal RGB-based methods in scenes with low-illumination or similar backgrounds. Most existing methods propose well-designed structures for cross-modal fusion in RGB-T crowd counting. However, these methods have difficulty in encoding cross-modal contextual semantic information in RGB-T image pairs. Considering the aforementioned problem, we propose a two-stream RGB-T crowd counting network called Multi-Attention Fusion Network (MAFNet), which aims to fully capture long-range contextual information from the RGB and thermal modalities based on the attention mechanism. Specifically, in the encoder part, a Multi-Attention Fusion (MAF) module is embedded into different stages of the two modality-specific branches for cross-modal fusion at the global level. In addition, a Multi-modal Multi-scale Aggregation (MMA) regression head is introduced to make full use of the multi-scale and contextual information across modalities to generate high-quality crowd density maps. Extensive experiments on two popular datasets show that the proposed MAFNet is effective for RGB-T crowd counting and achieves the state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2208.06761

Code (0)

등록된 구현이 없습니다.

Tasks

Crowd Counting

Similar Papers 제목 키워드 기반

Multi-scale Adaptive Fusion Network for Hyperspectral Image Denoising

2023-04-19 · Haodong Pan, Feng Gao, Junyu Dong, Qian Du

Removing the noise and improving the visual quality of hyperspectral images (HSIs) is challenging in academia and industry. Great efforts have been made to leverage local, global or spectral context information for HSI d…

DenoisingHyperspectral Image DenoisingImage Denoising

Multi-level Attention Fusion Network for Audio-visual Event Recognition

2021-06-12 · Mathilde Brousmiche, Jean Rouat, Stéphane Dupont

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level…

FusionCounting: Robust visible-infrared image fusion guided by crowd counting via multi-task learning

2025-08-28 · He Li, Xinyu Liu, Weihang Kong, Xingchen Zhang arxiv

Visible and infrared image fusion (VIF) is an important multimedia task in computer vision. Most VIF methods focus primarily on optimizing fused image quality. Recent studies have begun incorporating downstream tasks, su…

Semantic SegmentationMulti-Task LearningObject DetectionCrowd Counting

MAFNet:Multi-frequency Adaptive Fusion Network for Real-time Stereo Matching

2025-12-04 · Ao Xu, Rujin Zhao, Xiong Xu, Boceng Huang 외 arxiv

Existing stereo matching networks typically rely on either cost-volume construction based on 3D convolutions or deformation methods based on iterative optimization. The former incurs significant computational overhead du…

Disparity Estimation

Dual Path Multi-Scale Fusion Networks with Attention for Crowd Counting

2019-02-04 · Liang Zhu, Zhijian Zhao, Chao Lu, Yining Lin 외

The task of crowd counting in varying density scenes is an extremely difficult challenge due to large scale variations. In this paper, we propose a novel dual path multi-scale fusion network architecture with attention m…

Crowd Counting