Panoramic Vision Transformer for Saliency Detection in 360° Videos
360$^\circ$ video saliency detection is one of the challenging benchmarks for 360$^\circ$ video understanding since non-negligible distortion and discontinuity occur in the projection of any format of 360$^\circ$ videos, and capture-worthy viewpoint in the omnidirectional sphere is ambiguous by nature. We present a new framework named Panoramic Vision Transformer (PAVER). We design the encoder using Vision Transformer with deformable convolution, which enables us not only to plug pretrained models from normal videos into our architecture without additional modules or finetuning but also to perform geometric approximation only once, unlike previous deep CNN-based approaches. Thanks to its powerful encoder, PAVER can learn the saliency from three simple relative relations among local patch features, outperforming state-of-the-art models for the Wild360 benchmark by large margins without supervision or auxiliary information like class activation. We demonstrate the utility of our saliency prediction model with the omnidirectional video quality assessment task in VQA-ODV, where we consistently improve performance without any form of supervision, including head movement.
Code (1)
Tasks
Saliency DetectionSaliency PredictionVideo Quality AssessmentVideo Saliency DetectionVideo UnderstandingVisual Question Answering (VQA)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Panonut360: A Head and Eye Tracking Dataset for Panoramic Video
With the rapid development and widespread application of VR/AR technology, maximizing the quality of immersive panoramic video services that match users' personal preferences and habits has become a long-standing challen…
ASOD60K: An Audio-Induced Salient Object Detection Dataset for Panoramic Videos
Exploring to what humans pay attention in dynamic panoramic scenes is useful for many fundamental applications, including augmented reality (AR) in retail, AR-powered recruitment, and visual language navigation. With thi…
4kObjectobject-detectionObject Detection+2Automatic Salient Object Detection for Panoramic Images Using Region Growing and Fixation Prediction Model
Almost all previous works on saliency detection have been dedicated to conventional images, however, with the outbreak of panoramic images due to the rapid development of VR or AR technology, it is becoming more challeng…
Density Estimationobject-detectionObject DetectionRGB Salient Object Detection+2CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Video
Omnidirectional videos (ODVs) present distinct challenges for accurate audio-visual saliency prediction due to their immersive nature, which combines spatial audio with panoramic visuals to enhance the user experienc…
Saliency PredictionPanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation
Panoramic videos contain richer spatial information and have attracted tremendous amounts of attention due to their exceptional experience in some fields such as autonomous driving and virtual reality. However, existing …
Autonomous DrivingSegmentationSemantic SegmentationVideo Object Segmentation+2