paper-with-me

홈 › Papers

Generating image captions with external encyclopedic knowledge

2022-10-10 · Sofia Nikiforova, Tejaswini Deoskar, Denis Paperno, Yoad Winter

Accurately reporting what objects are depicted in an image is largely a solved problem in automatic caption generation. The next big challenge on the way to truly humanlike captioning is being able to incorporate the context of the image and related real world knowledge. We tackle this challenge by creating an end-to-end caption generation system that makes extensive use of image-specific encyclopedic data. Our approach includes a novel way of using image location to identify relevant open-domain facts in an external knowledge base, with their subsequent integration into the captioning pipeline at both the encoding and decoding stages. Our system is trained and tested on a new dataset with naturally produced knowledge-rich captions, and achieves significant improvements over multiple baselines. We empirically demonstrate that our approach is effective for generating contextualized captions with encyclopedic knowledge that is both factually accurate and relevant to the image.

📄 PDF Abstract BibTeX arXiv:2210.04806

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationImage CaptioningWorld Knowledge

Similar Papers 제목 키워드 기반

EchoSight: Advancing Visual-Language Models with Wiki Knowledge

2024-07-17 · Yibin Yan, Weidi Xie

Knowledge-based Visual Question Answering (KVQA) tasks require answering questions about images using extensive background knowledge. Despite significant advancements, generative models often struggle with these tasks du…

ArticlesQuestion AnsweringRAGRetrieval+3

ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering

2025-07-22 · Duong T. Tran, Trung-Kien Tran, Manfred Hauswirth, Danh Le Phuoc arxiv

In this paper, we propose a new dataset, ReasonVQA, for the Visual Question Answering (VQA) task. Our dataset is automatically integrated with structured encyclopedic knowledge and constructed using a low-cost framework,…

Visual Question Answering

Temporal Knowledge-Aware Image Captioning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Contextualized image captioning is a task that extends beyond generating a purely visual description of the image content and aims to produce a caption that is influenced by the context and informed by the real world kno…

Caption GenerationImage CaptioningWorld Knowledge

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

2026-03-05 · Shan Ning, Longtian Qiu, Xuming He arxiv

Knowledge-Based Visual Question Answering (KB-VQA) requires models to answer questions about an image by integrating external knowledge, posing significant challenges due to noisy retrieval and the structured, encycloped…

Visual Question AnsweringReinforcement LearningMultimodal ReasoningDomain Adaptation

Boost Image Captioning with Knowledge Reasoning

2020-11-02 · Feicheng Huang, Zhixin Li, Haiyang Wei, Canlong Zhang 외

Automatically generating a human-like description for a given image is a potential research in artificial intelligence, which has attracted a great of attention recently. Most of the existing attention methods explore th…

DecoderImage CaptioningSentence