paper-with-me

홈 › Papers

MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights

2026-01-08 · Jasmine Yang, Poppy Zhang, Shawndra Hill arxiv

We propose MLLM-VADStory, a novel domain knowledge-guided multimodal large language models (MLLM) framework to systematically quantify and generate insights for video ad storyline understanding at scale. The framework is centered on the core idea that ad narratives are structured by functional intent, with each scene unit performing a distinct communicative function, delivering product and brand-oriented information within seconds. MLLM-VADStory segments ads into functional units, classifies each unit's functionality using a novel advertising-specific functional role taxonomy, and then aggregates functional sequences across ads to recover data-driven storyline structures. Applying the framework to 50k social media video ads across four industry subverticals, we find that story-based creatives improve video retention, and we recommend top-performing story arcs to guide advertisers in creative design. Our framework demonstrates the value of using domain knowledge to guide MLLMs in generating scalable insights for video ad storylines, making it a versatile tool for understanding video creatives in general.

📄 PDF Abstract BibTeX arXiv:2601.07850

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

COSINT-Agent: A Knowledge-Driven Multimodal Agent for Chinese Open Source Intelligence

2025-03-05 · Wentao Li, Congcong Wang, Xiaoxiao Cui, Zhi Liu 외

Open Source Intelligence (OSINT) requires the integration and reasoning of diverse multimodal data, presenting significant challenges in deriving actionable insights. Traditional approaches, including multimodal large la…

Multimodal Reasoning

Empowering Source-Free Domain Adaptation with MLLM-driven Curriculum Learning

2024-05-28 · Dongjie Chen, Kartik Patwari, Zhengfeng Lai, Sen-ching Cheung 외

Source-Free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to a target domain using only unlabeled target data. Current SFDA methods face challenges in effectively leveraging pre-trained knowledge and …

Domain AdaptationInstruction FollowingSource-Free Domain AdaptationTransfer Learning+1

Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning

2026-01-23 · Qinglong Cao, Yuntian Chen, Chao Ma, Xiaokang Yang arxiv

Multimodal large language models (MLLMs) have shown remarkable capabilities in multimodal perception and understanding tasks. However, their effectiveness in specialized domains, such as remote sensing and medical imagin…

Domain Adaptation

InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning

2025-09-26 · Guanghao Zhu, Zhitian Hou, Zeyu Liu, Zhijie Sang 외 arxiv

Multimodal large language models (MLLMs) have shown remarkable potential in various domains, yet their application in the medical field is hindered by several challenges. General-purpose MLLMs often lack the specialized …

Visual Question AnsweringKnowledge DistillationContinual Pretraining

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

2026-07-27 · Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang 외 hf

We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: hol…