Zero-Shot Activity Recognition with Videos
In this paper, we examined the zero-shot activity recognition task with the usage of videos. We introduce an auto-encoder based model to construct a multimodal joint embedding space between the visual and textual manifolds. On the visual side, we used activity videos and a state-of-the-art 3D convolutional action recognition network to extract the features. On the textual side, we worked with GloVe word embeddings. The zero-shot recognition results are evaluated by top-n accuracy. Then, the manifold learning ability is measured by mean Nearest Neighbor Overlap. In the end, we provide an extensive discussion over the results and the future directions.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionActivity RecognitionWord EmbeddingsZero-Shot LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SEZ-HARN: Self-Explainable Zero-shot Human Activity Recognition Network
Human Activity Recognition (HAR), which uses data from Inertial Measurement Unit (IMU) sensors, has many practical applications in healthcare and assisted living environments. However, its use in real-world scenarios has…
Activity RecognitionHuman Activity RecognitionZero-Shot LearningZero-Shot Anticipation for Instructional Activities
How can we teach a robot to predict what will happen next for an activity it has never seen before? We address this problem of zero-shot anticipation by presenting a hierarchical model that generalizes instructional know…
Zero-Shot LearningZero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment
Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi samples for every target class, limiting their ability to recognize…
Human Activity RecognitionUnseen Action Recognition with Unpaired Adversarial Multimodal Learning
In this paper, we present a method to learn a joint multimodal representation space that allows for the recognition of unseen activities in videos. We compare the effect of placing various constraints on the embedding sp…
Action RecognitionGeneral ClassificationTemporal Action LocalizationVideo-Mined Task Graphs for Keystep Recognition in Instructional Videos
Procedural activity understanding requires perceiving human actions in terms of a broader task, where multiple keysteps are performed in sequence across a long video to reach a final goal state -- such as the steps of a …
Representation Learning