paper-with-me

Papers

Connecting Vision and Language with Localized Narratives

2019-12-06 · ECCV 2020 8 · Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut, Vittorio Ferrari

We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the region they are describing. Since the voice and the mouse pointer are synchronized, we can localize every single word in the description. This dense visual grounding takes the form of a mouse trace segment per word and is unique to our data. We annotated 849k images with Localized Narratives: the whole COCO, Flickr30k, and ADE20K datasets, and 671k images of Open Images, all of which we make publicly available. We provide an extensive analysis of these annotations showing they are diverse, accurate, and efficient to produce. We also demonstrate their utility on the application of controlled image captioning.

📄 PDF Abstract BibTeX arXiv:1912.03098

Code (1)

google/localized-narratives 공식 구현

Tasks

FormImage CaptioningImage GenerationVisual Grounding

Similar Papers 제목 키워드 기반

Connecting Vision and Language with Video Localized Narratives

2023-02-22 · CVPR 2023 1 · Paul Voigtlaender, Soravit Changpinyo, Jordi Pont-Tuset, Radu Soricut 외

We propose Video Localized Narratives, a new form of multimodal video annotations connecting vision and language. In the original Localized Narratives, annotators speak and move their mouse simultaneously on an image, th…

Question AnsweringVideo Narrative GroundingVideo Question Answering

MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives

2025-01-07 · Wisdom O. Ikezogwo, Kevin Zhang, Mehmet Saygin Seyfioglu, Fatemeh Ghezloo 외

We propose MedicalNarratives, a dataset curated from medical pedagogical videos similar in nature to data collected in Think-Aloud studies and inspired by Localized Narratives, which collects grounded image-text data by …

Articles

MaRU: A Manga Retrieval and Understanding System Connecting Vision and Language

2023-10-22 · Conghao Tom Shen, Violet Yao, Yixin Liu

Manga, a widely celebrated Japanese comic art form, is renowned for its diverse narratives and distinct artistic styles. However, the inherently visual and intricate structure of Manga, which comprises images housing mul…

Decoderobject-detectionObject DetectionRetrieval

GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives

2023-12-07 · Zuyao Chen, Jinlin Wu, Zhen Lei, Zhaoxiang Zhang 외

Training Scene Graph Generation (SGG) models with natural language captions has become increasingly popular due to the abundant, cost-effective, and open-world generalization supervision signals that natural language off…

Graph GenerationLanguage ModellingLarge Language ModelScene Graph Generation+1

Think Before You Act: A Two-Stage Framework for Mitigating Gender Bias Towards Vision-Language Tasks

2024-05-27 · Yunqi Zhang, Songda Li, Chunyuan Deng, Luyi Wang 외

Gender bias in vision-language models (VLMs) can reinforce harmful stereotypes and discrimination. In this paper, we focus on mitigating gender bias towards vision-language tasks. We identify object hallucination as the …

HallucinationObject Hallucination