paper-with-me

홈 › Papers

Panoramic Vision Transformer for Saliency Detection in 360° Videos

2022-09-19 · Heeseung Yun, Sehun Lee, Gunhee Kim

360$^\circ$ video saliency detection is one of the challenging benchmarks for 360$^\circ$ video understanding since non-negligible distortion and discontinuity occur in the projection of any format of 360$^\circ$ videos, and capture-worthy viewpoint in the omnidirectional sphere is ambiguous by nature. We present a new framework named Panoramic Vision Transformer (PAVER). We design the encoder using Vision Transformer with deformable convolution, which enables us not only to plug pretrained models from normal videos into our architecture without additional modules or finetuning but also to perform geometric approximation only once, unlike previous deep CNN-based approaches. Thanks to its powerful encoder, PAVER can learn the saliency from three simple relative relations among local patch features, outperforming state-of-the-art models for the Wild360 benchmark by large margins without supervision or auxiliary information like class activation. We demonstrate the utility of our saliency prediction model with the omnidirectional video quality assessment task in VQA-ODV, where we consistently improve performance without any form of supervision, including head movement.

📄 PDF Abstract BibTeX arXiv:2209.08956

Code (1)

hs-yn/paver 공식 구현 pytorch

Tasks

Saliency DetectionSaliency PredictionVideo Quality AssessmentVideo Saliency DetectionVideo UnderstandingVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Position-Wise Feed-Forward Layer 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Adam 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Panonut360: A Head and Eye Tracking Dataset for Panoramic Video

2024-03-26 · Yutong Xu, Junhao Du, Jiahe Wang, Yuwei Ning 외

With the rapid development and widespread application of VR/AR technology, maximizing the quality of immersive panoramic video services that match users' personal preferences and habits has become a long-standing challen…

ASOD60K: An Audio-Induced Salient Object Detection Dataset for Panoramic Videos

2021-07-24 · Yi Zhang

Exploring to what humans pay attention in dynamic panoramic scenes is useful for many fundamental applications, including augmented reality (AR) in retail, AR-powered recruitment, and visual language navigation. With thi…

4kObjectobject-detectionObject Detection+2

Automatic Salient Object Detection for Panoramic Images Using Region Growing and Fixation Prediction Model

2017-10-10 · Chunbiao Zhu, Kan Huang, Ge Li

Almost all previous works on saliency detection have been dedicated to conventional images, however, with the outbreak of panoramic images due to the rapid development of VR or AR technology, it is becoming more challeng…

Density Estimationobject-detectionObject DetectionRGB Salient Object Detection+2

CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Video

2025-01-01 · CVPR 2025 1 · Zhaolin Wan, Han Qin, Zhiyang Li, Xiaopeng Fan 외

Omnidirectional videos (ODVs) present distinct challenges for accurate audio-visual saliency prediction due to their immersive nature, which combines spatial audio with panoramic visuals to enhance the user experienc…

Saliency Prediction

PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation

2023-09-21 · Shilin Yan, Xiaohao Xu, Renrui Zhang, Lingyi Hong 외

Panoramic videos contain richer spatial information and have attracted tremendous amounts of attention due to their exceptional experience in some fields such as autonomous driving and virtual reality. However, existing …

Autonomous DrivingSegmentationSemantic SegmentationVideo Object Segmentation+2