A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Video
Dense object counting or crowd counting has come a long way thanks to the recent development in the vision community. However, indiscernible object counting, which aims to count the number of targets that are blended with respect to their surroundings, has been a challenge. Image-based object counting datasets have been the mainstream of the current publicly available datasets. Therefore, we propose a large-scale dataset called YoutubeFish-35, which contains a total of 35 sequences of high-definition videos with high frame-per-second and more than 150,000 annotated center points across a selected variety of scenes. For benchmarking purposes, we select three mainstream methods for dense object counting and carefully evaluate them on the newly collected dataset. We propose TransVidCount, a new strong baseline that combines density and regression branches along the temporal domain in a unified framework and can effectively tackle indiscernible object counting with state-of-the-art performance on YoutubeFish-35 dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingCrowd CountingObjectObject CountingSimilar Papers 제목 키워드 기반
RGB-D Indiscernible Object Counting in Underwater Scenes
Recently, indiscernible/camouflaged scene understanding has attracted lots of research attention in the vision community. We further advance the frontier of this field by systematically studying a new challenge named ind…
BenchmarkingDepth EstimationObjectObject Counting+1Counting Varying Density Crowds Through Density Guided Adaptive Selection CNN and Transformer Estimation
In real-world crowd counting applications, the crowd densities in an image vary greatly. When facing density variation, humans tend to locate and count the targets in low-density regions, and reason the number in high-de…
Crowd CountingFlow-Guided Transformer for Video Inpainting
We propose a flow-guided transformer, which innovatively leverage the motion discrepancy exposed by optical flows to instruct the attention retrieval in transformer for high fidelity video inpainting. More specially, we …
RetrievalVideo InpaintingPhysics-guided spatiotemporal neural models for fuel density prediction
This paper presents a physics-guided machine learning (PGML) framework for fuel density prediction, integrating physics constraints and domain knowledge into deep learning models to enhance model accuracy and stability. …
Extracting Temporal Event Relation with Syntax-guided Graph Transformer
Extracting temporal relations (e.g., before, after, and simultaneous) among events is crucial to natural language understanding. One of the key challenges of this problem is that when the events of interest are far away …
Dependency ParsingNatural Language UnderstandingRelationRelation Classification+3