paper-with-me

홈 › Papers

AQIFormer: A Transformer-Based Multi-View Architecture for Cross-City Air Quality Classification

2026-06-02 · Om Kathalkar, Nitin Nilesh, Sachin Chaudhari, Anoop Namboodiri arxiv

Air pollution represents one of the most critical environmental and public health challenges globally, with traditional sensor-based monitoring systems facing significant scalability and economic constraints. Image-based air quality estimation has emerged as a promising alternative, leveraging the visual characteristics of atmospheric pollutants in traffic scenes. However, existing methods suffer from limited cross-city generalization and inadequate exploitation of multi-view perspectives. We present AQIFormer, a novel transformer-based ensemble architecture that addresses these fundamental limitations through innovative dual-view integration, weather-aware attention mechanisms, and comprehensive multi-task learning. Our approach uniquely combines front and rear traffic imagery with meteorological parameters to achieve robust air quality classification across diverse urban environments. Extensive evaluation on a comprehensive dataset of 26,678 synchronized front-rear image pairs demonstrates good performance with 89.96% accuracy, representing a 14.96% improvement over state-of-the-art methods. Most importantly, our model maintains exceptional cross-city generalization capabilities, achieving 81.67% accuracy on an independent dataset collected in Nagpur, India with only 8.29% performance degradation using few-shot adaptation with minimal training samples.

📄 PDF Abstract BibTeX arXiv:2606.07648

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task Learning

Similar Papers 제목 키워드 기반

Cross-view Transformers for real-time Map-view Semantic Segmentation

2022-05-05 · CVPR 2022 1 · Brady Zhou, Philipp Krähenbühl

We present cross-view transformers, an efficient attention-based model for map-view semantic segmentation from multiple cameras. Our architecture implicitly learns a mapping from individual camera views into a canonical …

Bird's-Eye View Semantic SegmentationSegmentationSemantic Segmentation

XFMamba: Cross-Fusion Mamba for Multi-View Medical Image Classification

2025-03-04 · Xiaoyu Zheng, Xu Chen, Shaogang Gong, Xavier Griffin 외

Compared to single view medical image classification, using multiple views can significantly enhance predictive accuracy as it can account for the complementarity of each view while leveraging correlations between views.…

Classificationimage-classificationImage ClassificationMamba+1

Multiview Transformers for Video Recognition

2022-01-12 · CVPR 2022 1 · Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu 외

Video understanding requires reasoning at multiple spatiotemporal resolutions -- from short fine-grained motions to events taking place over longer durations. Although transformer architectures have recently advanced the…

Action ClassificationAction RecognitionVideo Understanding

Photorealistic Novel View Synthesis of Human Faces using Next-Scale Transformers

2026-08-24 · Federico Stella, Fei Jiang, Zhongshi Jiang, Zohar Barzelay 외 arxiv

Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving identity, fine appearance details, and geometric coherence is critical. W…

Novel View Synthesis

PFT-SSR: Parallax Fusion Transformer for Stereo Image Super-Resolution

2023-03-24 · Hansheng Guo, Juncheng Li, Guangwei Gao, Zhi Li 외

Stereo image super-resolution aims to boost the performance of image super-resolution by exploiting the supplementary information provided by binocular systems. Although previous methods have achieved promising results, …

Image Super-ResolutionStereo Image Super-ResolutionSuper-Resolution