Multimodal Prompt Learning for Product Title Generation with Extremely Limited Labels
Generating an informative and attractive title for the product is a crucial task for e-commerce. Most existing works follow the standard multimodal natural language generation approaches, e.g., image captioning, and employ the large scale of human-labelled datasets to train desirable models. However, for novel products, especially in a different domain, there are few existing labelled data. In this paper, we propose a prompt-based approach, i.e., the Multimodal Prompt Learning framework, to accurately and efficiently generate titles for novel products with limited labels. We observe that the core challenges of novel product title generation are the understanding of novel product characteristics and the generation of titles in a novel writing style. To this end, we build a set of multimodal prompts from different modalities to preserve the corresponding characteristics and writing styles of novel products. As a result, with extremely limited labels for training, the proposed method can retrieve the multimodal prompts to generate desirable titles for novel products. The experiments and analyses are conducted on five novel product categories under both the in-domain and out-of-domain experimental settings. The results show that, with only 1% of downstream labelled data for training, our proposed approach achieves the best few-shot results and even achieves competitive results with fully-supervised methods trained on 100% of training data; With the full labelled data for training, our method achieves state-of-the-art results.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningPrompt LearningText GenerationSimilar Papers 제목 키워드 기반
Enhancing E-commerce Product Title Translation with Retrieval-Augmented Generation and Large Language Models
E-commerce stores enable multilingual product discovery which require accurate product title translation. Multilingual large language models (LLMs) have shown promising capacity to perform machine translation tasks, and …
Machine TranslationRAGRetrievalRetrieval-augmented Generation+1Continuous Prompt Tuning Based Textual Entailment Model for E-commerce Entity Typing
The explosion of e-commerce has caused the need for processing and analysis of product titles, like entity typing in product titles. However, the rapid activity in e-commerce has led to the rapid emergence of new entitie…
Entity TypingNatural Language InferenceProduct Title Refinement via Multi-Modal Generative Adversarial Learning
Nowadays, an increasing number of customers are in favor of using E-commerce Apps to browse and purchase products. Since merchants are usually inclined to employ redundant and over-informative product titles to attract c…
AttributeGenerative Adversarial Networkreinforcement-learningReinforcement Learning+1E-MMAD: Multimodal Advertising Caption Generation Based on Structured Information
With multimodal tasks increasingly getting popular in recent years, datasets with large scale and reliable authenticity are in urgent demand. Therefore, we present an e-commercial multimodal advertising dataset, E-MMAD, …
Caption GenerationvalidVideo CaptioningMulti-Modal Generative Adversarial Network for Short Product Title Generation in Mobile E-Commerce
Nowadays, more and more customers browse and purchase products in favor of using mobile E-Commerce Apps such as Taobao and Amazon. Since merchants are usually inclined to describe redundant and over-informative product t…
AttributeGenerative Adversarial NetworkReinforcement Learning