paper-with-me

홈 › Papers

DensFiLM: Density-Conditioned Video Saliency for Crowd Scenes

2026-07-28 · Anis Ur Rahman arxiv

Video saliency models typically apply a single fixation strategy across crowd scenes, despite systematic changes in attention with crowd density. Sparse scenes encourage tracking individuals, whereas dense scenes shift attention toward collective motion and scene-level landmarks. We introduce DensFiLM, a density-conditioned video saliency model that inserts a lightweight Feature-wise Linear Modulation layer at the bottleneck of a Video Swin Transformer. A learned density embedding produces channel-wise scale and shift parameters, allowing the decoder to reconstruct saliency from features selected for each density regime. The module adds only ~100K parameters and can use either CrowdFix density labels or the model's own density prediction. On CrowdFix, DensFiLM achieves mean NSS 1.434 and CC 0.517 over four seeds, improving over ACLNet by 14.7% and 14.9%, respectively, while predicted-density conditioning matches oracle-label performance. Ablations show that explicit RAFT optical flow and larger temporal and social-force extensions provide no further improvement in this setting. In a centre-prior-subtraction diagnostic, density conditioning yields an NSS gain of 0.462 over the unconditioned backbone, compared with 0.124 under standard evaluation. These results show that lightweight bottleneck conditioning provides a more effective inductive bias than increasing model capacity for crowd-video saliency. Our code is available at https://github.com/aniskhan25/crowdfix-saliency.

📄 PDF Abstract BibTeX arXiv:2607.25465

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CrowdFix: An Eyetracking Dataset of Real Life Crowd Videos

2019-10-07 · Memoona Tahira, Sobas Mehboob, Anis U. Rahman, Omar Arif

Understanding human visual attention and saliency is an integral part of vision research. In this context, there is an ever-present need for fresh and diverse benchmark datasets, particularly for insight into special use…

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

2026-01-18 · Yaowu Fan, Jia Wan, Tao Han, Andy J. Ma 외 arxiv

Counting and tracking dense crowds in large-scale scenes is valuable yet challenging, while existing methods and datasets are largely limited to fixed cameras with small scene coverage. We introduce MovingDroneCrowd++, a…

Crowd Counting

Crowd Density Forecasting by Modeling Patch-based Dynamics

2019-11-22 · Hiroaki Minoura, Ryo Yonetani, Mai Nishimura, Yoshitaka Ushiku

Forecasting human activities observed in videos is a long-standing challenge in computer vision, which leads to various real-world applications such as mobile robots, autonomous driving, and assistive systems. In this wo…

Autonomous Driving

NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results

2026-04-16 · Andrey Moskalenko, Alexey Bryncev, Ivan Kosmynin, Kira Shilovskaya 외 arxiv

This paper presents an overview of the NTIRE 2026 Challenge on Video Saliency Prediction. The goal of the challenge participants was to develop automatic saliency map prediction methods for the provided video sequences. …

Saliency Prediction

Predicting video saliency using crowdsourced mouse-tracking data

2019-06-30 · Vitaliy Lyudvichenko, Dmitriy Vatolin

This paper presents a new way of getting high-quality saliency maps for video, using a cheaper alternative to eye-tracking data. We designed a mouse-contingent video viewing system which simulates the viewers' peripheral…

Position