paper-with-me

홈 › Papers

Look, Read and Feel: Benchmarking Ads Understanding with Multimodal Multitask Learning

2019-12-21 · Huaizheng Zhang, Yong Luo, Qiming Ai, Yonggang Wen

Given the massive market of advertising and the sharply increasing online multimedia content (such as videos), it is now fashionable to promote advertisements (ads) together with the multimedia content. It is exhausted to find relevant ads to match the provided content manually, and hence, some automatic advertising techniques are developed. Since ads are usually hard to understand only according to its visual appearance due to the contained visual metaphor, some other modalities, such as the contained texts, should be exploited for understanding. To further improve user experience, it is necessary to understand both the topic and sentiment of the ads. This motivates us to develop a novel deep multimodal multitask framework to integrate multiple modalities to achieve effective topic and sentiment prediction simultaneously for ads understanding. In particular, our model first extracts multimodal information from ads and learn high-level and comparable representations. The visual metaphor of the ad is decoded in an unsupervised manner. The obtained representations are then fed into the proposed hierarchical multimodal attention modules to learn task-specific representations for final prediction. A multitask loss function is also designed to train both the topic and sentiment prediction models jointly in an end-to-end manner. We conduct extensive experiments on the latest and large advertisement dataset and achieve state-of-the-art performance for both prediction tasks. The obtained results could be utilized as a benchmark for ads understanding.

📄 PDF Abstract BibTeX arXiv:1912.10248

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingPrediction

Similar Papers 제목 키워드 기반

Neural feels with neural fields: Visuo-tactile perception for in-hand manipulation

2023-12-20 · Sudharshan Suresh, Haozhi Qi, Tingfan Wu, Taosha Fan 외

To achieve human-level dexterity, robots must infer spatial awareness from multimodal sensing to reason over contact interactions. During in-hand manipulation of novel objects, such spatial awareness involves estimating …

Benchmarking

How you feelin'? Learning Emotions and Mental States in Movie Scenes

2023-04-12 · CVPR 2023 1 · Dhruv Srivastava, Aditya Kumar Singh, Makarand Tapaswi

Movie story analysis requires understanding characters' emotions and mental states. Towards this goal, we formulate emotion understanding as predicting a diverse and multi-label set of emotions at the level of a movie sc…

Emotion RecognitionMulti-Label Classification

Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

2026-06-25 · Jinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang Jiang arxiv

Unified multimodal models capable of both understanding and generation have achieved remarkable strides. However, despite their unified designs, existing evaluations typically assess understanding and generation capabili…

Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts

2026-08-31 · Jaehee Kim, Ji Hoon Chung, Seoyoon Park, Unsol Kim 외 arxiv

Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this re…

Question Answering

MMTF-DES: A Fusion of Multimodal Transformer Models for Desire, Emotion, and Sentiment Analysis of Social Media Data

2023-10-22 · Abdul Aziz, Nihad Karim Chowdhury, Muhammad Ashad Kabir, Abu Nowshed Chy 외

Desire is a set of human aspirations and wishes that comprise verbal and cognitive aspects that drive human feelings and behaviors, distinguishing humans from other animals. Understanding human desire has the potential t…

Emotional IntelligenceEmotion RecognitionSentiment Analysis