paper-with-me

홈 › Papers

Attention Filtering for Multi-person Spatiotemporal Action Detection on Deep Two-Stream CNN Architectures

2019-07-21 · João Antunes, Pedro Abreu, Alexandre Bernardino, Asim Smailagic, Daniel Siewiorek

Action detection and recognition tasks have been the target of much focus in the computer vision community due to their many applications, namely, security, robotics and recommendation systems. Recently, datasets like AVA, provide multi-person, multi-label, spatiotemporal action detection and recognition challenges. Being unable to discern which portions of the input to use for classification is a limitation of two-stream CNN approaches, once the vision task involves several people with several labels. We address this limitation and improve the state-of-the-art performance of two-stream CNNs. In this paper we present four contributions: our fovea attention filtering that highlights targets for classification without discarding background; a generalized binary loss function designed for the AVA dataset; miniAVA, a partition of AVA that maintains temporal continuity and class distribution with only one tenth of the dataset size; and ablation studies on alternative attention filters. Our method, using fovea attention filtering and our generalized binary loss, achieves a relative video mAP improvement of 20% over the two-stream baseline in AVA, and is competitive with the state-of-the-art in the UCF101-24. We also show a relative video mAP improvement of 12.6% when using our generalized binary loss over the standard sum-of-sigmoids.

📄 PDF Abstract BibTeX arXiv:1907.12919

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionGeneral ClassificationRecommendation Systems

Similar Papers 제목 키워드 기반

Spatiotemporal-Enhanced Network for Click-Through Rate Prediction in Location-based Services

2022-09-20 · Shaochuan Lin, Yicong Yu, Xiyu Ji, Taotao Zhou 외

In Location-Based Services(LBS), user behavior naturally has a strong dependence on the spatiotemporal information, i.e., in different geographical locations and at different times, user click behavior will change signif…

AttributeClick-Through Rate Prediction

Spatiotemporal Filtering for Event-Based Action Recognition

2019-03-17 · Rohan Ghosh, Anupam Gupta, Andrei Nakagawa, Alcimar Soares 외

In this paper, we address the challenging problem of action recognition, using event-based cameras. To recognise most gestural actions, often higher temporal precision is required for sampling visual information. Actions…

Action RecognitionTemporal Action Localization

Snipper: A Spatiotemporal Transformer for Simultaneous Multi-Person 3D Pose Estimation Tracking and Forecasting on a Video Snippet

2022-07-09 · Shihao Zou, Yuanlu Xu, Chao Li, Lingni Ma 외

Multi-person pose understanding from RGB videos involves three complex tasks: pose estimation, tracking and motion forecasting. Intuitively, accurate multi-person pose estimation facilitates robust tracking, and robust t…

3D Pose EstimationMotion ForecastingMulti-Person Pose EstimationPose Estimation

Deep Spatiotemporal Clutter Filtering of Transthoracic Echocardiographic Images: Leveraging Contextual Attention and Residual Learning

2024-01-23 · Mahdi Tabassian, Somayeh Akbari, Sandro Queirós, Jan D'hooge

This study presents a deep convolutional autoencoder network for filtering reverberation clutter from transthoracic echocardiographic (TTE) image sequences. Given the spatiotemporal nature of this type of clutter, the fi…

Synergetic Reconstruction from 2D Pose and 3D Motion for Wide-Space Multi-Person Video Motion Capture in the Wild

2020-01-16 · Takuya Ohashi, Yosuke Ikegami, Yoshihiko Nakamura

Although many studies have investigated markerless motion capture, the technology has not been applied to real sports or concerts. In this paper, we propose a markerless motion capture method with spatiotemporal accuracy…

3D ReconstructionMarkerless Motion Capture