paper-with-me

홈 › Papers

CTR-Driven Advertising Image Generation with Multimodal Large Language Models

2025-02-05 · Xingye Chen, Wei Feng, Zhenbang Du, Weizhen Wang, Yanyin Chen, Haohan Wang, Linkai Liu, Yaoyu Li, Jinyuan Zhao, Yu Li, Zheng Zhang, Jingjing Lv, Junjie Shen, Zhangang Lin, Jingping Shao, Yuanjie Shao, Xinge You, Changxin Gao, Nong Sang

In web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily focus on the aesthetic quality, which may fail to achieve satisfactory online performance. To address this limitation, we explore the use of Multimodal Large Language Models (MLLMs) for generating advertising images by optimizing for Click-Through Rate (CTR) as the primary objective. Firstly, we build targeted pre-training tasks, and leverage a large-scale e-commerce multimodal dataset to equip MLLMs with initial capabilities for advertising image generation tasks. To further improve the CTR of generated images, we propose a novel reward model to fine-tune pre-trained MLLMs through Reinforcement Learning (RL), which can jointly utilize multimodal features and accurately reflect user click preferences. Meanwhile, a product-centric preference optimization strategy is developed to ensure that the generated background content aligns with the product characteristics after fine-tuning, enhancing the overall relevance and effectiveness of the advertising images. Extensive experiments have demonstrated that our method achieves state-of-the-art performance in both online and offline metrics. Our code and pre-trained models are publicly available at: https://github.com/Chenguoz/CAIG.

📄 PDF Abstract BibTeX arXiv:2502.06823

Code (1)

chenguoz/caig 공식 구현 pytorch

Tasks

Image GenerationReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Agentic Multimodal AI for Hyperpersonalized B2B and B2C Advertising in Competitive Markets: An AI-Driven Competitive Advertising Framework

2025-04-01 · Sakhinana Sagar Srinivas, Akash Das, Shivam Gupta, Venkataramana Runkana

The growing use of foundation models (FMs) in real-world applications demands adaptive, reliable, and efficient strategies for dynamic markets. In the chemical industry, AI-discovered materials drive innovation, but comm…

Decision MakingIn-Context LearningMarketingMultimodal Reasoning+3

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

2026-05-12 · Yexing Xu, Wei Feng, Shen Zhang, Haohan Wang 외 arxiv

Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models driven by click-through-rate (CTR) to controllably create attractive image …

Text Generation

One Size, Many Fits: Aligning Diverse Group-Wise Click Preferences in Large-Scale Advertising Image Generation

2026-02-02 · Shuo Lu, Haohan Wang, Wei Feng, Weizhen Wang 외 arxiv

Advertising image generation has increasingly focused on online metrics like Click-Through Rate (CTR), yet existing approaches adopt a ``one-size-fits-all" strategy that optimizes for overall CTR while neglecting prefere…

Image Generation

KD-CVG: A Knowledge-Driven Approach for Creative Video Generation

2026-04-23 · Linkai Liu, Wei Feng, Xi Zhao, Shen Zhang 외 arxiv

Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a significant focus of recent research. However, while CG has advanced c…

Reinforcement LearningVideo Generation

E-MMAD: Multimodal Advertising Caption Generation Based on Structured Information

2021-11-16 · ACL ARR November 2021 11 · Anonymous

With multimodal tasks increasingly getting popular in recent years, datasets with large scale and reliable authenticity are in urgent demand. Therefore, we present an e-commercial multimodal advertising dataset, E-MMAD, …

Caption GenerationvalidVideo Captioning