paper-with-me

Papers

Spatio-Temporal Event Segmentation and Localization for Wildlife Extended Videos

2020-05-05 · Ramy Mounir, Roman Gula, Jörn Theuerkauf, Sudeep Sarkar

Using offline training schemes, researchers have tackled the event segmentation problem by providing full or weak-supervision through manually annotated labels or self-supervised epoch-based training. Most works consider videos that are at most 10's of minutes long. We present a self-supervised perceptual prediction framework capable of temporal event segmentation by building stable representations of objects over time and demonstrate it on long videos, spanning several days. The approach is deceptively simple but quite effective. We rely on predictions of high-level features computed by a standard deep learning backbone. For prediction, we use an LSTM, augmented with an attention mechanism, trained in a self-supervised manner using the prediction error. The self-learned attention maps effectively localize and track the event-related objects in each frame. The proposed approach does not require labels. It requires only a single pass through the video, with no separate training set. Given the lack of datasets of very long videos, we demonstrate our method on video from 10 days (254 hours) of continuous wildlife monitoring data that we had collected with required permissions. We find that the approach is robust to various environmental conditions such as day/night conditions, rain, sharp shadows, and windy conditions. For the task of temporally locating events, we had an 80% recall rate at 20% false-positive rate for frame-level segmentation. At the activity level, we had an 80% activity recall rate for one false activity detection every 50 minutes. We will make the dataset, which is the first of its kind, and the code available to the research community.

📄 PDF Abstract BibTeX arXiv:2005.02463

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionContinual LearningEvent SegmentationSegmentation

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Trimmed Action Recognition, Dense-Captioning Events in Videos, and Spatio-temporal Action Localization with Focus on ActivityNet Challenge 2019

2019-06-14 · Zhaofan Qiu, Dong Li, Yehao Li, Qi Cai 외

This notebook paper presents an overview and comparative analysis of our systems designed for the following three tasks in ActivityNet Challenge 2019: trimmed action recognition, dense-captioning events in videos, and sp…

Action LocalizationAction RecognitionDense CaptioningSpatio-Temporal Action Localization+1

Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical Flow

2025-03-10 · CVPR 2025 1 · Hanyu Zhou, Haonan Wang, Haoyue Liu, Yuxing Duan 외

High-dynamic scene optical flow is a challenging task, which suffers spatial blur and temporal discontinuous motion due to large displacement in frame imaging, thus deteriorating the spatiotemporal feature of optical flo…

Optical Flow Estimation

SIS-Challenge: Event-based Spatio-temporal Instance Segmentation Challenge at the CVPR 2025 Event-based Vision Workshop

2025-08-18 · Friedhelm Hamann, Emil Mededovic, Fabian Gülhan, Yuli Wu 외 arxiv

We present an overview of the Spatio-temporal Instance Segmentation (SIS) challenge held in conjunction with the CVPR 2025 Event-based Vision Workshop. The task is to predict accurate pixel-level segmentation masks of de…

Instance SegmentationEvent-based vision

Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

2026-06-12 · Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida, Kim Sung-Bin 외 arxiv

Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track sou…

YH Technologies at ActivityNet Challenge 2018

2018-06-29 · Ting Yao, Xue Li

This notebook paper presents an overview and comparative analysis of our systems designed for the following five tasks in ActivityNet Challenge 2018: temporal action proposals, temporal action localization, dense-caption…

Action LocalizationAction RecognitionDense CaptioningSpatio-Temporal Action Localization+1