Predicting Online Video Advertising Effects with Multimodal Deep Learning
With expansion of the video advertising market, research to predict the effects of video advertising is getting more attention. Although effect prediction of image advertising has been explored a lot, prediction for video advertising is still challenging with seldom research. In this research, we propose a method for predicting the click through rate (CTR) of video advertisements and analyzing the factors that determine the CTR. In this paper, we demonstrate an optimized framework for accurately predicting the effects by taking advantage of the multimodal nature of online video advertisements including video, text, and metadata features. In particular, the two types of metadata, i.e., categorical and continuous, are properly separated and normalized. To avoid overfitting, which is crucial in our task because the training data are not very rich, additional regularization layers are inserted. Experimental results show that our approach can achieve a correlation coefficient as high as 0.695, which is a significant improvement from the baseline (0.487).
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningMultimodal Deep LearningSimilar Papers 제목 키워드 기반
AgenticGen: Reward-Guided Agentic Video Generation for Advertising
Advertising video generation is not only a video synthesis task, but also a product-conditioned reasoning problem whose success is measured by online business metrics. Recent video foundation models can generate realisti…
Video GenerationContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising
Contextual advertising serves ads that are aligned to the content that the user is viewing. The rapid growth of video content on social platforms and streaming services, along with privacy concerns, has increased the nee…
RetrievalText to Video RetrievalVideo RetrievalAlignment Helps Make the Most of Multimodal Data
When studying political communication, combining the information from text, audio, and video signals promises to reflect the richness of human communication more comprehensively than confining it to individual modalities…
On the Effects of Video Grounding on Language Models
Transformer-based models trained on text and vision modalities try to improve the performance on multimodal downstream tasks or tackle the problem Transformer-based models trained on text and vision modalities try to imp…
Image CaptioningQuestion AnsweringVideo GroundingVisual Question Answering+1Online Causal Inference for Advertising in Real-Time Bidding Auctions
Real-time bidding (RTB) systems, which utilize auctions to allocate user impressions to competing advertisers, continue to enjoy success in digital advertising. Assessing the effectiveness of such advertising remains a c…
Causal InferenceExperimental DesignThompson Sampling