paper-with-me

Papers

Going Deeper with Semantics: Video Activity Interpretation using Semantic Contextualization

2017-08-11 · Sathyanarayanan N. Aakur, Fillipe DM de Souza, Sudeep Sarkar

A deeper understanding of video activities extends beyond recognition of underlying concepts such as actions and objects: constructing deep semantic representations requires reasoning about the semantic relationships among these concepts, often beyond what is directly observed in the data. To this end, we propose an energy minimization framework that leverages large-scale commonsense knowledge bases, such as ConceptNet, to provide contextual cues to establish semantic relationships among entities directly hypothesized from video signal. We mathematically express this using the language of Grenander's canonical pattern generator theory. We show that the use of prior encoded commonsense knowledge alleviate the need for large annotated training datasets and help tackle imbalance in training through prior knowledge. Using three different publicly available datasets - Charades, Microsoft Visual Description Corpus and Breakfast Actions datasets, we show that the proposed model can generate video interpretations whose quality is better than those reported by state-of-the-art approaches, which have substantial training needs. Through extensive experiments, we show that the use of commonsense knowledge from ConceptNet allows the proposed approach to handle various challenges such as training data imbalance, weak features, and complex semantic relationships and visual scenes.

📄 PDF Abstract BibTeX arXiv:1708.03725

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Hybrid Graph Network for Complex Activity Detection in Video

2023-10-26 · Salman Khan, Izzeddin Teeti, Andrew Bradley, Mohamed Elhoseiny 외

Interpretation and understanding of video presents a challenging computer vision task in numerous fields - e.g. autonomous driving and sports analytics. Existing approaches to interpreting the actions taking place within…

Action DetectionActivity DetectionAllAutonomous Driving+2

Long Activity Video Understanding using Functional Object-Oriented Network

2018-07-03 · Ahmad Babaeian Jelodar, David Paulius, Yu Sun

Video understanding is one of the most challenging topics in computer vision. In this paper, a four-stage video understanding pipeline is presented to simultaneously recognize all atomic actions and the single on-going a…

ObjectVideo Understanding

Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval

2026-01-19 · Zequn Xie, Boyun Zhang, Yuxiao Lin, Tao Jin arxiv

Video-text retrieval (VTR) aims to locate relevant videos using natural language queries. Current methods, often based on pre-trained models like CLIP, are hindered by video's inherent redundancy and their reliance on co…

Natural Language QueriesVideo-Text Retrieval

SBGAR: Semantics Based Group Activity Recognition

2017-10-01 · ICCV 2017 10 · Xin Li, Mooi Choo Chuah

Activity recognition has become an important function in many emerging computer vision applications e.g. automatic video surveillance system, human-computer interaction application, and video recommendation system, etc. …

Activity RecognitionGroup Activity Recognition

Early Recognition of Human Activities from First-Person Videos Using Onset Representations

2014-06-20 · M. S. Ryoo, Thomas J. Fuchs, Lu Xia, J. K. Aggarwal 외

In this paper, we propose a methodology for early recognition of human activities from videos taken with a first-person viewpoint. Early recognition, which is also known as activity prediction, is an ability to infer an …

Activity PredictionPerson RecognitionTime SeriesTime Series Analysis