Connecting Vision and Language with Video Localized Narratives
We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move their mouse simultaneously on an image, thus grounding each word with a mouse trace segment. However, this is challenging on a video. Our new protocol empowers annotators to tell the story of a video with Localized Narratives, capturing even complex events involving multiple actors interacting with each other and with several passive objects. We annotated 20k videos of the OVIS, UVO, and Oops datasets, totalling 1.7M words. Based on this data, we also construct new benchmarks for the video narrative grounding and video question answering tasks, and provide reference results from strong baseline models. Our annotations are available at https://google.github.io/video-localized-narratives/.
Code (1)
Tasks
Question AnsweringVideo Narrative GroundingVideo Question AnsweringSimilar Papers 제목 키워드 기반
Connecting Vision and Language with Localized Narratives
We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the regio…
FormImage CaptioningImage GenerationVisual GroundingMedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
We propose MedicalNarratives, a dataset curated from medical pedagogical videos similar in nature to data collected in Think-Aloud studies and inspired by Localized Narratives, which collects grounded image-text data by …
ArticlesTraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, strug…
Trajectory PredictionScene UnderstandingLogical ReasoningBridging Vision and Language: Modeling Causality and Temporality in Video Narratives
Video captioning is a critical task in the field of multimodal machine learning, aiming to generate descriptive and coherent textual narratives for video content. While large vision-language models (LVLMs) have shown sig…
DescriptiveLanguage ModelingLanguage ModellingVideo CaptioningMaRU: A Manga Retrieval and Understanding System Connecting Vision and Language
Manga, a widely celebrated Japanese comic art form, is renowned for its diverse narratives and distinct artistic styles. However, the inherently visual and intricate structure of Manga, which comprises images housing mul…
Decoderobject-detectionObject DetectionRetrieval