paper-with-me

홈 › Papers

Grounding Action Descriptions in Videos

2013-01-01 · TACL 2013 1 · Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, Manfred Pinkal

Recent work has shown that the integration of visual information into text-based models can substantially improve model predictions, but so far only visual information extracted from static images has been used. In this paper, we consider the problem of grounding sentences describing actions in visual information extracted from videos. We present a general purpose corpus that aligns high quality videos with multiple natural language descriptions of the actions portrayed in the videos, together with an annotation of how similar the action descriptions are to each other. Experimental results demonstrate that a text-based model of similarity between actions improves substantially when combined with visual information from videos depicting the described actions.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Textual SimilarityVideo Understanding

Similar Papers 제목 키워드 기반

MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions

2021-12-01 · CVPR 2022 1 · Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron 외

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at asse…

Moment RetrievalNatural Language Moment Retrieval

What When and Where? Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions

2024-01-01 · CVPR 2024 1 · Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann 외

Spatio-temporal grounding describes the task of localizing events in space and time e.g. in video data based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bo…

Representation Learning

Character Grounding and Re-Identification in Story of Videos and Text Descriptions

2020-08-01 · ECCV 2020 8 · Youngjae Yu, Jongseok Kim, Heeseung Yun, Jiwan Chung 외

We address character grounding and re-identification in multiple story-based videos like movies and associated text descriptions. In order to solve these related tasks in a mutually rewarding way, we propose a model name…

Gender Prediction

Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

2019-04-08 · WS 2019 6 · Peratham Wiriyathammabhum, Abhinav Shrivastava, Vlad I. Morariu, Larry S. Davis

This paper presents a new task, the grounding of spatio-temporal identifying descriptions in videos. Previous work suggests potential bias in existing datasets and emphasizes the need for a new data creation schema to be…

What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions

2023-03-29 · Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann 외

Spatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bou…

Representation LearningSpatio-Temporal Video Grounding