paper-with-me

홈 › Papers

Can You Spot the Semantic Predicate in this Video?

2018-08-01 · COLING 2018 8 · Christopher Reale, Claire Bonial, Heesung Kwon, Clare Voss

We propose a method to improve human activity recognition in video by leveraging semantic information about the target activities from an expert-defined linguistic resource, VerbNet. Our hypothesis is that activities that share similar event semantics, as defined by the semantic predicates of VerbNet, will be more likely to share some visual components. We use a deep convolutional neural network approach as a baseline and incorporate linguistic information from VerbNet through multi-task learning. We present results of experiments showing the added information has negligible impact on recognition performance. We discuss how this may be because the lexical semantic information defined by VerbNet is generally not visually salient given the video processing approach used here, and how we may handle this in future approaches.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Activity RecognitionHuman Activity RecognitionMulti-Task Learning

Similar Papers 제목 키워드 기반

SPOT! Revisiting Video-Language Models for Event Understanding

2023-11-21 · Gengyuan Zhang, Jinhe Bi, Jindong Gu, Yanyu Chen 외

Understanding videos is an important research topic for multimodal learning. Leveraging large-scale datasets of web-crawled video-text pairs as weak supervision has become a pre-training paradigm for learning joint repre…

AttributeVideo Understanding

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

2026-08-26 · Eunjee Choi, JungHoon Sung, Seongwhan Cho, Chu Xin 외 arxiv

Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, provid…

Sign Language RecognitionRepresentation Learning

A Context-Aware Loss Function for Action Spotting in Soccer Videos

2019-12-03 · CVPR 2020 6 · Anthony Cioppa, Adrien Deliège, Silvio Giancola, Bernard Ghanem 외

In video understanding, action spotting consists in temporally localizing human-induced events annotated with single timestamps. In this paper, we propose a novel loss function that specifically considers the temporal co…

Action SpottingVideo Understanding

Grounded Human-Object Interaction Hotspots from Video (Extended Abstract)

2019-06-03 · Tushar Nagarajan, Christoph Feichtenhofer, Kristen Grauman

Learning how to interact with objects is an important step towards embodied visual intelligence, but existing techniques suffer from heavy supervision or sensing requirements. We propose an approach to learn human-object…

Human-Object Interaction DetectionObjectSemantic Segmentation

Cross-Modal Graph with Meta Concepts for Video Captioning

2021-08-14 · Hao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan Miao

Video captioning targets interpreting the complex visual contents as text descriptions, which requires the model to fully understand video scenes including objects and their interactions. Prevailing methods adopt off-the…

object-detectionObject DetectionVideo Captioning