Video In Sentences Out
We present a system that produces sentential descriptions of video: who did what to whom, and where and how they did it. Action class is rendered as a verb, participant objects as noun phrases, properties of those objects as adjectival modifiers in those noun phrases, spatial relations between those participants as prepositional phrases, and characteristics of the event as prepositional-phrase adjuncts and adverbial modifiers. Extracting the information needed to render these linguistic entities requires an approach to event recognition that recovers object tracks, the trackto-role assignments, and changing body posture.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Weakly-Supervised Temporal Article Grounding
Given a long untrimmed video and natural language queries, video grounding (VG) aims to temporally localize the semantically-aligned video segments. Almost all existing VG work holds two simple but unrealistic assumption…
AllArticlesNatural Language QueriesSentence+1Open-book Video Captioning with Retrieve-Copy-Generate Network
Due to the rapid emergence of short videos and the requirement for content understanding and creation, the video captioning task has received increasing attention in recent years. In this paper, we convert traditional vi…
DecoderRetrievalVideo CaptioningPseudo-labeling with Keyword Refining for Few-Supervised Video Captioning
Video captioning generate a sentence that describes the video content. Existing methods always require a number of captions (\eg, 10 or 20) per video to train the model, which is quite costly. In this work, we explore th…
Video CaptioningSyntax Customized Video Captioning by Imitating Exemplar Sentences
Enhancing the diversity of sentences to describe video contents is an important problem arising in recent video captioning research. In this paper, we explore this problem from a novel perspective of customizing video ca…
DecoderDiversitySentencevalid+1Video Captioning Using Weak Annotation
Video captioning has shown impressive progress in recent years. One key reason of the performance improvements made by existing methods lie in massive paired video-sentence data, but collecting such strong annotation, i.…
SentenceVideo CaptioningVisual Reasoning