paper-with-me

홈 › Papers

Exploiting Image–Text Synergy for Contextual Image Captioning

2021-04-01 · EACL (LANTERN) 2021 4 · Sreyasi Nag Chowdhury, Rajarshi Bhowmik, Hareesh Ravi, Gerard de Melo, Simon Razniewski, Gerhard Weikum

Modern web content - news articles, blog posts, educational resources, marketing brochures - is predominantly multimodal. A notable trait is the inclusion of media such as images placed at meaningful locations within a textual narrative. Most often, such images are accompanied by captions - either factual or stylistic (humorous, metaphorical, etc.) - making the narrative more engaging to the reader. While standalone image captioning has been extensively studied, captioning an image based on external knowledge such as its surrounding text remains under-explored. In this paper, we study this new task: given an image and an associated unstructured knowledge snippet, the goal is to generate a contextual caption for the image.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesImage CaptioningMarketing

Similar Papers 제목 키워드 기반

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

2026-05-26 · Xiang Fang, Wanlong Fang, Changshuo Wang arxiv

Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness …

Visual Question AnsweringAutonomous DrivingImage Captioning

Draft-and-Revise: Effective Image Generation with Contextual RQ-Transformer

2022-06-09 · Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho 외

Although autoregressive models have achieved promising results on image generation, their unidirectional generation process prevents the resultant images from fully reflecting global contexts. To address the issue, we pr…

Conditional Image GenerationImage GenerationText-to-Image Generation

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models

2026-04-13 · Songlin Yang, Xianghao Kong, Anyi Rao arxiv

Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of vision models. In practice, however, this synergy remains elusive: UMMs fa…

Text-to-Image GenerationText Generation

Context-based Deep Learning Architecture with Optimal Integration Layer for Image Parsing

2022-04-13 · Ranju Mandal, Basim Azam, Brijesh Verma

Deep learning models have been efficient lately on image parsing tasks. However, deep learning models are not fully capable of exploiting visual and contextual information simultaneously. The proposed three-layer context…

Deep Learning

Prompted Contextual Transformer for Incomplete-View CT Reconstruction

2023-12-13 · Chenglong Ma, Zilong Li, Junjun He, Junping Zhang 외

Incomplete-view computed tomography (CT) can shorten the data acquisition time and allow scanning of large objects, including sparse-view and limited-angle scenarios, each with various settings, such as different view nu…

Computed Tomography (CT)CT Reconstruction