paper-with-me

Papers

Connecting Vision and Language with Video Localized Narratives

2023-02-22 · CVPR 2023 1 · Paul Voigtlaender, Soravit Changpinyo, Jordi Pont-Tuset, Radu Soricut, Vittorio Ferrari

We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move their mouse simultaneously on an image, thus grounding each word with a mouse trace segment. However, this is challenging on a video. Our new protocol empowers annotators to tell the story of a video with Localized Narratives, capturing even complex events involving multiple actors interacting with each other and with several passive objects. We annotated 20k videos of the OVIS, UVO, and Oops datasets, totalling 1.7M words. Based on this data, we also construct new benchmarks for the video narrative grounding and video question answering tasks, and provide reference results from strong baseline models. Our annotations are available at https://google.github.io/video-localized-narratives/.

📄 PDF Abstract BibTeX arXiv:2302.11217

Code (1)

google/video-localized-narratives 공식 구현

Tasks

Question AnsweringVideo Narrative GroundingVideo Question Answering

Similar Papers 제목 키워드 기반

Connecting Vision and Language with Localized Narratives

2019-12-06 · ECCV 2020 8 · Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut 외

We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the regio…

FormImage CaptioningImage GenerationVisual Grounding

MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives

2025-01-07 · Wisdom O. Ikezogwo, Kevin Zhang, Mehmet Saygin Seyfioglu, Fatemeh Ghezloo 외

We propose MedicalNarratives, a dataset curated from medical pedagogical videos similar in nature to data collected in Think-Aloud studies and inspired by Localized Narratives, which collects grounded image-text data by …

Articles

TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

2026-02-23 · Fan Yang, Shurong Zheng, Hongyin Zhao, Yufei Zhan 외 arxiv

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, strug…

Trajectory PredictionScene UnderstandingLogical Reasoning

Bridging Vision and Language: Modeling Causality and Temporality in Video Narratives

2024-12-14 · Ji-jun Park, Soo-joon Choi

Video captioning is a critical task in the field of multimodal machine learning, aiming to generate descriptive and coherent textual narratives for video content. While large vision-language models (LVLMs) have shown sig…

DescriptiveLanguage ModelingLanguage ModellingVideo Captioning

MaRU: A Manga Retrieval and Understanding System Connecting Vision and Language

2023-10-22 · Conghao Tom Shen, Violet Yao, Yixin Liu

Manga, a widely celebrated Japanese comic art form, is renowned for its diverse narratives and distinct artistic styles. However, the inherently visual and intricate structure of Manga, which comprises images housing mul…

Decoderobject-detectionObject DetectionRetrieval