Exploiting Image–Text Synergy for Contextual Image Captioning
Modern web content - news articles, blog posts, educational resources, marketing brochures - is predominantly multimodal. A notable trait is the inclusion of media such as images placed at meaningful locations within a textual narrative. Most often, such images are accompanied by captions - either factual or stylistic (humorous, metaphorical, etc.) - making the narrative more engaging to the reader. While standalone image captioning has been extensively studied, captioning an image based on external knowledge such as its surrounding text remains under-explored. In this paper, we study this new task: given an image and an associated unstructured knowledge snippet, the goal is to generate a contextual caption for the image.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesImage CaptioningMarketingSimilar Papers 제목 키워드 기반
Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization
Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness …
Visual Question AnsweringAutonomous DrivingImage CaptioningDraft-and-Revise: Effective Image Generation with Contextual RQ-Transformer
Although autoregressive models have achieved promising results on image generation, their unidirectional generation process prevents the resultant images from fully reflecting global contexts. To address the issue, we pr…
Conditional Image GenerationImage GenerationText-to-Image GenerationPseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models
Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of vision models. In practice, however, this synergy remains elusive: UMMs fa…
Text-to-Image GenerationText GenerationContext-based Deep Learning Architecture with Optimal Integration Layer for Image Parsing
Deep learning models have been efficient lately on image parsing tasks. However, deep learning models are not fully capable of exploiting visual and contextual information simultaneously. The proposed three-layer context…
Deep LearningPrompted Contextual Transformer for Incomplete-View CT Reconstruction
Incomplete-view computed tomography (CT) can shorten the data acquisition time and allow scanning of large objects, including sparse-view and limited-angle scenarios, each with various settings, such as different view nu…
Computed Tomography (CT)CT Reconstruction