paper-with-me

Papers

Multimodal Crowd Counting with Pix2Pix GANs

2024-01-15 · Muhammad Asif Khan, Hamid Menouar, Ridha Hamila

Most state-of-the-art crowd counting methods use color (RGB) images to learn the density map of the crowd. However, these methods often struggle to achieve higher accuracy in densely crowded scenes with poor illumination. Recently, some studies have reported improvement in the accuracy of crowd counting models using a combination of RGB and thermal images. Although multimodal data can lead to better predictions, multimodal data might not be always available beforehand. In this paper, we propose the use of generative adversarial networks (GANs) to automatically generate thermal infrared (TIR) images from color (RGB) images and use both to train crowd counting models to achieve higher accuracy. We use a Pix2Pix GAN network first to translate RGB images to TIR images. Our experiments on several state-of-the-art crowd counting models and benchmark crowd datasets report significant improvement in accuracy.

📄 PDF Abstract BibTeX arXiv:2401.07591

Code (0)

등록된 구현이 없습니다.

Tasks

Crowd Counting

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
PatchGAN 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
Batch Normalization 설명 없음
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Why Existing Multimodal Crowd Counting Datasets Can Lead to Unfulfilled Expectations in Real-World Applications

2023-04-13 · Martin Thißen, Elke Hergenröther

More information leads to better decisions and predictions, right? Confirming this hypothesis, several studies concluded that the simultaneous use of optical and thermal images leads to better predictions in crowd counti…

Crowd Counting

Crowd Counting in Harsh Weather using Image Denoising with Pix2Pix GANs

2023-10-11 · Muhammad Asif Khan, Hamid Menouar, Ridha Hamila

Visual crowd counting estimates the density of the crowd using deep learning models such as convolution neural networks (CNNs). The performance of the model heavily relies on the quality of the training data that constit…

Crowd CountingDenoisingGenerative Adversarial NetworkImage Denoising

Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting

2020-12-08 · CVPR 2021 1 · Lingbo Liu, Jiaqi Chen, Hefeng Wu, Guanbin Li 외

Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot we…

Crowd CountingRepresentation Learning

Free Lunch Enhancements for Multi-modal Crowd Counting

2025-01-01 · CVPR 2025 1 · Haoliang Meng, Xiaopeng Hong, Zhengqin Lai, Miao Shang

This paper addresses multi-modal crowd counting with a novel `free lunch' training enhancement strategy that requires no additional data, parameters, or increased inference complexity. First, we introduce a cross-mod…

cross-modal alignmentCrowd Counting

A Benchmark for Semi-supervised Multi-modal Crowd Counting

2026-06-02 · Haoliang Meng, Xiaopeng Hong, Yabin Wang, Wangmeng Zuo arxiv

This paper constructs the first benchmark on semi-supervised multi-modal crowd counting. To lay the foundation for this unexplored task, we first formulate the semi-supervised multi-modal setting and a standardized proto…

Crowd Counting