paper-with-me

홈 › Papers

Intra- and Inter-Action Understanding via Temporal Action Parsing

2020-05-20 · CVPR 2020 6 · Dian Shao, Yue Zhao, Bo Dai, Dahua Lin

Current methods for action recognition primarily rely on deep convolutional networks to derive feature embeddings of visual and motion features. While these methods have demonstrated remarkable performance on standard benchmarks, we are still in need of a better understanding as to how the videos, in particular their internal structures, relate to high-level semantics, which may lead to benefits in multiple aspects, e.g. interpretable predictions and even new methods that can take the recognition performances to a next level. Towards this goal, we construct TAPOS, a new dataset developed on sport videos with manual annotations of sub-actions, and conduct a study on temporal action parsing on top. Our study shows that a sport activity usually consists of multiple sub-actions and that the awareness of such temporal structures is beneficial to action recognition. We also investigate a number of temporal parsing methods, and thereon devise an improved method that is capable of mining sub-actions from training data without knowing the labels of them. On the constructed TAPOS, the proposed method is shown to reveal intra-action information, i.e. how action instances are made of sub-actions, and inter-action information, i.e. one specific sub-action may commonly appear in various actions.

📄 PDF Abstract BibTeX arXiv:2005.10229

Code (0)

등록된 구현이 없습니다.

Tasks

Action ParsingAction RecognitionAction Understanding

Similar Papers 제목 키워드 기반

Instrument-tissue Interaction Detection Framework for Surgical Video Understanding

2024-03-30 · Wenjun Lin, Yan Hu, Huazhu Fu, Mingming Yang 외

Instrument-tissue interaction detection task, which helps understand surgical activities, is vital for constructing computer-assisted surgery systems but with many challenges. Firstly, most models represent instrument-ti…

Video Understanding

Spatio-Temporal Graph Transformer Networks for Pedestrian Trajectory Prediction

2020-05-18 · ECCV 2020 8 · Cunjun Yu, Xiao Ma, Jiawei Ren, Haiyu Zhao 외

Understanding crowd motion dynamics is critical to real-world applications, e.g., surveillance systems and autonomous driving. This is challenging because it requires effectively modeling the socially aware crowd spatial…

Autonomous DrivingPedestrian Trajectory PredictionPredictionTrajectory Prediction

Snippet-Aware Transformer With Multiple Action Elements for Skeleton-Based Action Segmentation

2024-05-06 · IEEE Transactions on Neural Networks and Learning Systems 2024 5 · Haoyu Ji, Bowen Chen, Wenze Huang, Weihong Ren 외

The skeleton-based temporal action segmentation (STAS) aims to densely segment and classify human actions within lengthy untrimmed skeletal motion sequences. Current methods primarily rely on graph convolutional networks…

Action SegmentationSkeleton Based Action SegmentationTemporal Action SegmentationVideo Understanding

Multi-Modal Interaction Graph Convolutional Network for Temporal Language Localization in Videos

2021-10-12 · Zongmeng Zhang, Xianjing Han, Xuemeng Song, Yan Yan 외

This paper focuses on tackling the problem of temporal language localization in videos, which aims to identify the start and end points of a moment described by a natural language sentence in an untrimmed video. However,…

Semantic correspondenceSemantic SimilaritySemantic Textual SimilaritySentence

FusionTransNet for Smart Urban Mobility: Spatiotemporal Traffic Forecasting Through Multimodal Network Integration

2024-05-09 · Binwu Wang, Yan Leng, Guang Wang, Yang Wang

This study develops FusionTransNet, a framework designed for Origin-Destination (OD) flow predictions within smart and multimodal urban transportation systems. Urban transportation complexity arises from the spatiotempor…

Decoder