paper-with-me

홈 › Papers

Multimodal Image Captioning for Marketing Analysis

2018-02-06 · Philipp Harzig, Stephan Brehm, Rainer Lienhart, Carolin Kaiser, René Schallner

Automatically captioning images with natural language sentences is an important research topic. State of the art models are able to produce human-like sentences. These models typically describe the depicted scene as a whole and do not target specific objects of interest or emotional relationships between these objects in the image. However, marketing companies require to describe these important attributes of a given scene. In our case, objects of interest are consumer goods, which are usually identifiable by a product logo and are associated with certain brands. From a marketing point of view, it is desirable to also evaluate the emotional context of a trademarked product, i.e., whether it appears in a positive or a negative connotation. We address the problem of finding brands in images and deriving corresponding captions by introducing a modified image captioning network. We also add a third output modality, which simultaneously produces real-valued image ratings. Our network is trained using a classification-aware loss function in order to stimulate the generation of sentences with an emphasis on words identifying the brand of a product. We evaluate our model on a dataset of images depicting interactions between humans and branded products. The introduced network improves mean class accuracy by 24.5 percent. Thanks to adding the third output modality, it also considerably improves the quality of generated captions for images depicting branded products.

📄 PDF Abstract BibTeX arXiv:1802.01958

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningMarketing

Similar Papers 제목 키워드 기반

Exploiting Image–Text Synergy for Contextual Image Captioning

2021-04-01 · EACL (LANTERN) 2021 4 · Sreyasi Nag Chowdhury, Rajarshi Bhowmik, Hareesh Ravi, Gerard de Melo 외

Modern web content - news articles, blog posts, educational resources, marketing brochures - is predominantly multimodal. A notable trait is the inclusion of media such as images placed at meaningful locations within a t…

ArticlesImage CaptioningMarketing

Social Media Ready Caption Generation for Brands

2024-01-03 · Himanshu Maheshwari, Koustava Goswami, Apoorv Saxena, Balaji Vasan Srinivasan

Social media advertisements are key for brand marketing, aiming to attract consumers with captivating captions and pictures or logos. While previous research has focused on generating captions for general images, incorpo…

Caption GenerationImage CaptioningMarketing

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis

2024-12-04 · Davide Bucciarelli, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 외

The task of image captioning demands an algorithm to generate natural language descriptions of visual inputs. Recent advancements have seen a convergence between image captioning research and the development of Large Lan…

Image CaptioningImage DescriptionPrompt Learning

Towards Better Graph Representation: Two-Branch Collaborative Graph Neural Networks for Multimodal Marketing Intention Detection

2020-05-13 · Lu Zhang, Jian Zhang, Zhibin Li, Jingsong Xu

Inspired by the fact that spreading and collecting information through the Internet becomes the norm, more and more people choose to post for-profit contents (images and texts) in social networks. Due to the difficulty o…

Graph ClassificationMarketing

QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models

2026-01-10 · Jiale Wang, Gee Wah Ng, Lee Onn Mak, Randall Cher 외 arxiv

This paper introduces QCaption, a novel video captioning and Q&A pipeline that enhances video analytics by fusing three models: key frame extraction, a Large Multimodal Model (LMM) for image-text analysis, and a Large La…

Video Captioning