paper-with-me

Papers

E-MMAD: Multimodal Advertising Caption Generation Based on Structured Information

2021-11-16 · ACL ARR November 2021 11 · Anonymous

With multimodal tasks increasingly getting popular in recent years, datasets with large scale and reliable authenticity are in urgent demand. Therefore, we present an e-commercial multimodal advertising dataset, E-MMAD, which contains 120 thousand valid data elaborately picked out from 1.3 million real product examples in both Chinese and English. Noticeably, it is one of the largest video captioning datasets in this field, in which each example has its product video (around 30 seconds), title, caption and structured information table that is observed to play a vital role in practice. We also introduce a fresh task for vision-language research based on E-MMAD: e-commercial multimodal advertising generation, which requires to use aforementioned product multimodal information to generate textual advertisement. Accordingly, we propose a baseline method on the strength of structured information reasoning to solve the demand in reality on this dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationvalidVideo Captioning

Similar Papers 제목 키워드 기반

Attract me to Buy: Advertisement Copywriting Generation with Multimodal Multi-structured Information

2022-05-07 · Zhipeng Zhang, Xinglin Hou, Kai Niu, Zhongzhen Huang 외

Recently, online shopping has gradually become a common way of shopping for people all over the world. Wonderful merchandise advertisements often attract more people to buy. These advertisements properly integrate multim…

Text GenerationVideo Captioning

MMaDA: Multimodal Large Diffusion Language Models

2025-05-21 · Ling Yang, Ye Tian, Bowen Li, Xinchen Zhang 외

We introduce MMaDA, a novel class of multimodal diffusion foundation models designed to achieve superior performance across diverse domains such as textual reasoning, multimodal understanding, and text-to-image generatio…

Image GenerationReinforcement Learning (RL)Text to Image GenerationText-to-Image Generation

Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models

2025-07-22 · Tz-Ying Wu, Tahani Trigui, Sharath Nittur Sridhar, Anand Bodas 외 arxiv

In this paper, we introduce VideoNarrator, a novel training-free pipeline designed to generate dense video captions that offer a structured snapshot of video content. These captions offer detailed narrations with precise…

Video Question AnsweringVideo Summarization

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

2025-11-12 · Ye Tian, Ling Yang, Jiongfan Yang, Anran Wang 외 arxiv

While thinking-aware generation aims to improve performance on complex tasks, we identify a critical failure mode where existing sequential, autoregressive approaches can paradoxically degrade performance due to error pr…

Reinforcement Learning

Agentic Multimodal AI for Hyperpersonalized B2B and B2C Advertising in Competitive Markets: An AI-Driven Competitive Advertising Framework

2025-04-01 · Sakhinana Sagar Srinivas, Akash Das, Shivam Gupta, Venkataramana Runkana

The growing use of foundation models (FMs) in real-world applications demands adaptive, reliable, and efficient strategies for dynamic markets. In the chemical industry, AI-discovered materials drive innovation, but comm…

Decision MakingIn-Context LearningMarketingMultimodal Reasoning+3