paper-with-me

홈 › Papers

Invariant recognition drives neural representations of action sequences

2017-04-20

Recognizing the actions of others from visual stimuli is a crucial aspect of human visual perception that allows individuals to respond to social cues. Humans are able to identify similar behaviors and discriminate between distinct actions despite transformations, like changes in viewpoint or actor, that substantially alter the visual appearance of a scene. This ability to generalize across complex transformations is a hallmark of human visual intelligence. Advances in understanding motion perception at the neural level have not always translated in precise accounts of the computational principles underlying what representation our visual cortex evolved or learned to compute. Here we test the hypothesis that invariant action discrimination might fill this gap. Recently, the study of artificial systems for static object perception has produced models, CNNs, that achieve human level performance in complex discriminative tasks. Within this class of models, architectures that better support invariant object recognition also produce image representations that match those implied by human and primate neural data. However, whether these models produce representations of action sequences that support recognition across complex transformations and closely follow neural representations remains unknown. Here we show that spatiotemporal CNNs appropriately categorize video stimuli into actions, and that deliberate model modifications that improve performance on an invariant action recognition task lead to data representations that better match human neural recordings. Our results support our hypothesis that performance on invariant discrimination dictates the neural representations of actions computed by human visual cortex.

📄 PDF Abstract BibTeX arXiv:1606.04698

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionObject Recognition

Similar Papers 제목 키워드 기반

Spatial-Temporal Alignment Network for Action Recognition

2023-08-19 · Jinhui Ye, Junwei Liang

This paper studies introducing viewpoint invariant feature representations in existing action recognition architecture. Despite significant progress in action recognition, efficiently handling geometric variations in lar…

Action Recognition

Recognizing Actions in Videos from Unseen Viewpoints

2021-03-30 · CVPR 2021 1 · AJ Piergiovanni, Michael S. Ryoo

Standard methods for video recognition use large CNNs designed to capture spatio-temporal data. However, training these models requires a large amount of labeled training data, containing a wide variety of actions, scene…

Action ClassificationAction RecognitionVideo Recognition

Spatial-Temporal Alignment Network for Action Recognition and Detection

2020-12-04 · Junwei Liang, Liangliang Cao, Xuehan Xiong, Ting Yu 외

This paper studies how to introduce viewpoint-invariant feature representations that can help action recognition and detection. Although we have witnessed great progress of action recognition in the past decade, it remai…

Action DetectionAction Recognition

ActAR: Actor-Driven Pose Embeddings for Video Action Recognition

2022-04-19 · Soufiane Lamghari, Guillaume-Alexandre Bilodeau, Nicolas Saunier

Human action recognition (HAR) in videos is one of the core tasks of video understanding. Based on video sequences, the goal is to recognize actions performed by humans. While HAR has received much attention in the visib…

Action RecognitionOptical Flow EstimationTemporal Action LocalizationVideo Understanding

Unsupervised Learning of View-invariant Action Representations

2018-09-06 · NeurIPS 2018 12 · Junnan Li, Yongkang Wong, Qi Zhao, Mohan S. Kankanhalli

The recent success in human action recognition with deep learning methods mostly adopt the supervised learning paradigm, which requires significant amount of manually labeled data to achieve good performance. However, la…

Action RecognitionRepresentation LearningTemporal Action Localization