Papers Video Saliency Prediction
“Video Saliency Prediction” 태그가 달린 논문 29편 · 필터 해제
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
Video saliency prediction is crucial for downstream applications, such as video compression and human-computer interaction. With the flourishing of multimodal learning, researchers started to explore multimodal video sal…
DenoisingImage GenerationPredictionSaliency Prediction+2DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
Audio-visual saliency prediction aims to mimic human visual attention by identifying salient regions in videos through the integration of both visual and auditory information. Although visual-only approaches have signifi…
Computational EfficiencySaliency PredictionVideo Saliency PredictionMinimalistic Video Saliency Prediction via Efficient Decoder & Spatio Temporal Action Cues
This paper introduces ViNet-S, a 36MB model based on the ViNet architecture with a U-Net design, featuring a lightweight decoder that significantly reduces model size and parameters without compromising performance. Addi…
Action ClassificationAction LocalizationDecoderSaliency Prediction+3Relevance-guided Audio Visual Fusion for Video Saliency Prediction
Audio data, often synchronized with video frames, plays a crucial role in guiding the audience's visual attention. Incorporating audio information into video saliency prediction tasks can enhance the prediction of human …
PredictionSaliency PredictionVideo Saliency PredictionAIM 2024 Challenge on Video Saliency Prediction: Methods and Results
This paper reviews the Challenge on Video Saliency Prediction at AIM 2024. The goal of the participants was to develop a method for predicting accurate saliency maps for the provided set of video sequences. Saliency maps…
Saliency DetectionSaliency PredictionVideo CompressionVideo Saliency Detection+1CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
Video saliency prediction aims to identify the regions in a video that attract human attention and gaze, driven by bottom-up features from the video and top-down processes like memory and cognition. Among these top-down …
Language ModellingLarge Language ModelMultimodal Large Language ModelPrediction+2SalFoM: Dynamic Saliency Prediction with Video Foundation Models
Recent advancements in video saliency prediction (VSP) have shown promising performance compared to the human visual system, whose emulation is the primary goal of VSP. However, current state-of-the-art models employ spa…
DecoderPredictionSaliency PredictionVideo Saliency PredictionTransformer-based Video Saliency Prediction with High Temporal Dimension Decoding
In recent years, finding an effective and efficient strategy for exploiting spatial and temporal information has been a hot research topic in video saliency prediction (VSP). With the emergence of spatio-temporal transfo…
DecoderSaliency PredictionVideo Saliency PredictionUniST: Towards Unifying Saliency Transformer for Video Saliency Prediction and Detection
Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceiving dynamic scenes. While many approaches have crafte…
Decoderobject-detectionObject DetectionPrediction+4Spherical Vision Transformer for 360-degree Video Saliency Prediction
The growing interest in omnidirectional videos (ODVs) that capture the full field-of-view (FOV) has gained 360-degree saliency prediction importance in computer vision. However, predicting where humans look in 360-degree…
PredictionSaliency PredictionVideo Saliency PredictionVideo UnderstandingCASP-Net: Rethinking Video Saliency Prediction from an Audio-VisualConsistency Perceptual Perspective
Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods a…
DecoderSaliency PredictionVideo Saliency PredictionTinyHD: Efficient Video Saliency Prediction with Heterogeneous Decoders using Hierarchical Maps Distillation
Video saliency prediction has recently attracted attention of the research community, as it is an upstream task for several practical applications. However, current solutions are particularly computationally demanding, e…
Knowledge DistillationPredictionSaliency PredictionVideo Saliency PredictionCASP-Net: Rethinking Video Saliency Prediction From an Audio-Visual Consistency Perceptual Perspective
Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP metho…
DecoderSaliency PredictionVideo Saliency PredictionGASP: Gated Attention For Saliency Prediction
Saliency prediction refers to the computational task of modeling overt attention. Social cues greatly influence our attention, consequently altering our eye movements and behavior. To emphasize the efficacy of such featu…
PredictionSaliency PredictionVideo Saliency DetectionVideo Saliency PredictionSpatio-Temporal Self-Attention Network for Video Saliency Prediction
3D convolutional neural networks have achieved promising results for video tasks in computer vision, including video saliency prediction that is explored in this paper. However, 3D convolution encodes visual representati…
PredictionSaliency PredictionVideo Saliency PredictionNoise-Aware Video Saliency Prediction
We tackle the problem of predicting saliency maps for videos of dynamic scenes. We note that the accuracy of the maps reconstructed from the gaze data of a fixed number of observers varies with the frame, as it depends o…
PredictionSaliency PredictionVideo Saliency PredictionViNet: Pushing the limits of Visual Modality for Audio-Visual Saliency Prediction
We propose the ViNet architecture for audio-visual saliency prediction. ViNet is a fully convolutional encoder-decoder architecture. The encoder uses visual features from a network trained for action recognition, and the…
Action RecognitionDecoderPredictionSaliency Prediction+2Hierarchical Domain-Adapted Feature Learning for Video Saliency Prediction
In this work, we propose a 3D fully convolutional architecture for video saliency prediction that employs hierarchical supervision on intermediate maps (referred to as conspicuity maps) generated using features extracted…
Domain AdaptationSaliency DetectionSaliency PredictionUnsupervised Domain Adaptation+2Video Saliency Prediction Using Enhanced Spatiotemporal Alignment Network
Due to a variety of motions across different frames, it is highly challenging to learn an effective spatiotemporal representation for accurate video saliency prediction (VSP). To address this issue, we develop an effecti…
PredictionSaliency PredictionVideo Saliency DetectionVideo Saliency PredictionSimple vs complex temporal recurrences for video saliency prediction
This paper investigates modifying an existing neural network architecture for static saliency prediction using two types of recurrences that integrate information from the temporal domain. The first modification is the a…
PredictionSaliency PredictionVideo Saliency DetectionVideo Saliency Prediction