Person Count Localization in Videos From Noisy Foreground and Detections
This paper formulates and presents a solution to a new problem called person count localization. Given a video of a crowded scene, our goal is to output for each frame a set of: 1) Detections optimally covering both isolated individuals and cluttered groups of people; and 2) Counts of people inside these detections. This problem is a middle-ground between frame-level person counting, which does not localize counts, and person detection aimed at perfectly localizing people with count-one detections. Our problem formulation is important for a wide range of domains, where people appear frequently under severe occlusion within a crowd. As these crowds are often visually distinct from the rest of the scene, they can be viewed as ``visual phrases'' whose spatially tight localization and count assignment could facilitate higher-level video understanding. For count localization, we specify a novel framework of iterative error-driven revisions of a flow graph derived from noisy input of people detections and foreground segmentation. Each iteration creates and solves an integer program for count localization based on iterative revisions of the flow graph. The graph revisions are based on detected violations of basic integrity constraints. They in turn trigger learned modifications to the graph aimed at reducing noise in input features. For evaluation, we introduce a new metric that measures both count precision and localization of our approach on American football and pedestrian videos.
Code (0)
등록된 구현이 없습니다.
Tasks
Foreground SegmentationHuman DetectionVideo UnderstandingSimilar Papers 제목 키워드 기반
Dance Dance Generation: Motion Transfer for Internet Videos
This work presents computational methods for transferring body movements from one person to another with videos collected in the wild. Specifically, we train a personalized model on a single video from the Internet which…
PAMI-AD: An Activity Detector Exploiting Part-attention and Motion Information in Surveillance Videos
Activity detection in surveillance videos is a challenging task caused by small objects, complex activity categories, its untrimmed nature, etc. Existing methods are generally limited in performance due to inaccurate pro…
Action DetectionActivity DetectionMulti-Object TrackingObject TrackingAdaptive Proposal Generation Network for Temporal Sentence Localization in Videos
We address the problem of temporal sentence localization in videos (TSLV). Traditional methods follow a top-down framework which localizes the target segment with pre-defined segment proposals. Although they have achieve…
SentenceForeground Clustering for Joint Segmentation and Localization in Videos and Images
This paper presents a novel framework in which video/image segmentation and localization are cast into a single optimization problem that integrates information from low level appearance cues with that of high level loca…
ClusteringImage SegmentationObjectObject Discovery+3Online Localization and Prediction of Actions and Interactions
This paper proposes a person-centric and online approach to the challenging problem of localization and prediction of actions and interactions in videos. Typically, localization or recognition is performed in an offline …
Pose EstimationPredictionSuperpixels