paper-with-me

홈 › Papers

Natural Language Descriptions for Human Activities in Video Streams

2017-09-01 · WS 2017 9 · Nouf Alharbi, Yoshihiko Gotoh

There has been continuous growth in the volume and ubiquity of video material. It has become essential to define video semantics in order to aid the searchability and retrieval of this data. We present a framework that produces textual descriptions of video, based on the visual semantic content. Detected action classes rendered as verbs, participant objects converted to noun phrases, visual properties of detected objects rendered as adjectives and spatial relations between objects rendered as prepositions. Further, in cases of zero-shot action recognition, a language model is used to infer a missing verb, aided by the detection of objects and scene settings. These extracted features are converted into textual descriptions using a template-based approach. The proposed video descriptions framework evaluated on the NLDHA dataset using ROUGE scores and human judgment evaluation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionLanguage ModelingLanguage ModellingRetrievalTemporal Action LocalizationText GenerationZero-Shot Action Recognition

Similar Papers 제목 키워드 기반

Temporal Modular Networks for Retrieving Complex Compositional Activities in Videos

2018-09-01 · ECCV 2018 9 · Bingbin Liu, Serena Yeung, Edward Chou, De-An Huang 외

A major challenge in computer vision is scaling activity understanding to the long tail of complex activities without requiring collecting large quantities of data for new actions. The task of video retrieval using natur…

RetrievalVideo Retrieval

Natural Language Descriptions of Human Activities Scenes: Corpus Generation and Analysis

2016-08-01 · WS 2016 8 · Nouf Alharbi, Yoshihiko Gotoh
Action ClassificationObject RecognitionText GenerationVideo Description

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

2023-11-30 · CVPR 2024 1 · Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani 외

We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g.,…

Video Understanding

Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them

2024-08-21 · Harish Haresamudram, Apoorva Beedu, Mashfiqui Rabbi, Sankalita Saha 외

Cross-modal contrastive pre-training between natural language and other modalities, e.g., vision and audio, has demonstrated astonishing performance and effectiveness across a diverse variety of tasks and domains. In thi…

Activity RecognitionCross-Modal RetrievalHuman Activity Recognition

OSVidCap: A Framework for the Simultaneous Recognition and Description of Concurrent Actions in Videos in an Open-Set Scenario

2021-09-29 · IEEE Access 2021 9 · Andrei De Souza Inácio, Matheus Gutoski, André Eugênio Lazzaretti, Heitor Silvério Lopes

Automatically understanding and describing the visual content of videos in natural language is a challenging task in computer vision. Existing approaches are often designed to describe single events in a closed-set setti…

DecoderOpen Set Video CaptioningVideo Captioning