paper-with-me

Papers

Image Captioning at Will: A Versatile Scheme for Effectively Injecting Sentiments into Image Descriptions

2018-01-30 · Quanzeng You, Hailin Jin, Jiebo Luo

Automatic image captioning has recently approached human-level performance due to the latest advances in computer vision and natural language understanding. However, most of the current models can only generate plain factual descriptions about the content of a given image. However, for human beings, image caption writing is quite flexible and diverse, where additional language dimensions, such as emotion, humor and language styles, are often incorporated to produce diverse, emotional, or appealing captions. In particular, we are interested in generating sentiment-conveying image descriptions, which has received little attention. The main challenge is how to effectively inject sentiments into the generated captions without altering the semantic matching between the visual content and the generated descriptions. In this work, we propose two different models, which employ different schemes for injecting sentiments into image captions. Compared with the few existing approaches, the proposed models are much simpler and yet more effective. The experimental results show that our model outperform the state-of-the-art models in generating sentimental (i.e., sentiment-bearing) image captions. In addition, we can also easily manipulate the model by assigning different sentiments to the testing image to generate captions with the corresponding sentiments.

📄 PDF Abstract BibTeX arXiv:1801.10121

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningNatural Language Understanding

Similar Papers 제목 키워드 기반

Image Change Captioning by Learning From an Auxiliary Task

2021-06-19 · CVPR 2021 1 · Mehrdad Hosseinzadeh, Yang Wang

We tackle the challenging task of image change captioning. The goal is to describe the subtle difference between two very similar images by generating a sentence caption. While the recent methods mainly focus on prop…

Image RetrievalMulti-Task LearningRetrievalSentence

Protect, Show, Attend and Tell: Empowering Image Captioning Models with Ownership Protection

2020-08-25 · Jian Han Lim, Chee Seng Chan, Kam Woh Ng, Lixin Fan 외

By and large, existing Intellectual Property (IP) protection on deep neural networks typically i) focus on image classification task only, and ii) follow a standard digital watermarking framework that was conventionally …

Image Captioningimage-classificationImage Classification

OmniCaptioner: One Captioner to Rule Them All

2025-04-09 · Yiting Lu, Jiakang Yuan, Zhen Li, Shitian Zhao 외

We propose OmniCaptioner, a versatile visual captioning framework for generating fine-grained textual descriptions across a wide variety of visual domains. Unlike prior methods limited to specific image types (e.g., natu…

AllImage CaptioningImage GenerationText to Image Generation+2

Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval

2024-04-11 · CVPR 2024 1 · Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 외

There has been significant attention to the research on dense video captioning, which aims to automatically localize and caption all events within untrimmed video. Several studies introduce methods by designing dense vid…

DecoderDense Video CaptioningRetrievalText Matching+1

OneDiff: A Generalist Model for Image Difference Captioning

2024-07-08 · Erdong Hu, Longteng Guo, Tongtian Yue, Zijia Zhao 외

In computer vision, Image Difference Captioning (IDC) is crucial for accurately describing variations between closely related images. Traditional IDC methods often rely on specialist models, which restrict their applicab…

Language ModellingmodelMulti-Task Learning