paper-with-me

Papers

Resource-Efficient Multiview Perception: Integrating Semantic Masking with Masked Autoencoders

2024-10-07 · Kosta Dakic, Kanchana Thilakarathna, Rodrigo N. Calheiros, Teng Joon Lim

Multiview systems have become a key technology in modern computer vision, offering advanced capabilities in scene understanding and analysis. However, these systems face critical challenges in bandwidth limitations and computational constraints, particularly for resource-limited camera nodes like drones. This paper presents a novel approach for communication-efficient distributed multiview detection and tracking using masked autoencoders (MAEs). We introduce a semantic-guided masking strategy that leverages pre-trained segmentation models and a tunable power function to prioritize informative image regions. This approach, combined with an MAE, reduces communication overhead while preserving essential visual information. We evaluate our method on both virtual and real-world multiview datasets, demonstrating comparable performance in terms of detection and tracking performance metrics compared to state-of-the-art techniques, even at high masking ratios. Our selective masking algorithm outperforms random masking, maintaining higher accuracy and precision as the masking ratio increases. Furthermore, our approach achieves a significant reduction in transmission data volume compared to baseline methods, thereby balancing multiview tracking performance with communication efficiency.

📄 PDF Abstract BibTeX arXiv:2410.04817

Code (0)

등록된 구현이 없습니다.

Tasks

Multiview DetectionScene Understanding

Methods 이 논문이 사용한 방법론

MAE 설명 없음

Similar Papers 제목 키워드 기반

End-to-End 3-D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration

2025-12-26 · Zhenwei Yang, Yibo Ai, Weidong Zhang arxiv

Multiview cooperative perception and multimodal fusion are essential for reliable 3-D spatiotemporal understanding in autonomous driving, especially in cases with occlusions, limited viewpoints, and communication delays …

Autonomous Driving

MSFormer: A Skeleton-multiview Fusion Method For Tooth Instance Segmentation

2023-10-23 · Yuan Li, Huan Liu, Yubo Tao, Xiangyang He 외

Recently, deep learning-based tooth segmentation methods have been limited by the expensive and time-consuming processes of data collection and labeling. Achieving high-precision segmentation with limited datasets is cri…

Contrastive LearningInstance SegmentationSegmentationSemantic Segmentation

What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

2026-06-24 · Salini Yadav, Taveena Lotey, Pravendra Singh, Partha Pratim Roy arxiv

Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging due to the low signal-to-noise ratio, non-stationarity, and limited …

Representation LearningContrastive LearningGraph Learning

SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

2022-06-21 · Gang Li, Heliang Zheng, Daqing Liu, Chaoyue Wang 외

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (M…

Language ModelingLanguage ModellingMasked Language ModelingSemantic Segmentation

Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video Recognition

2023-08-22 · ICCV 2023 1 · Qitong Wang, Long Zhao, Liangzhe Yuan, Ting Liu 외

We are concerned with a challenging scenario in unpaired multiview video learning. In this case, the model aims to learn comprehensive multiview representations while the cross-view semantic information exhibits variatio…

Multiview LearningVideo Recognition