paper-with-me

Papers

A Case Study of Deep Learning-Based Multi-Modal Methods for Labeling the Presence of Questionable Content in Movie Trailers

2021-09-01 · RANLP 2021 9 · Mahsa Shafaei, Christos Smailis, Ioannis Kakadiaris, Thamar Solorio

In this work, we explore different approaches to combine modalities for the problem of automated age-suitability rating of movie trailers. First, we introduce a new dataset containing videos of movie trailers in English downloaded from IMDB and YouTube, along with their corresponding age-suitability rating labels. Secondly, we propose a multi-modal deep learning pipeline addressing the movie trailer age suitability rating problem. This is the first attempt to combine video, audio, and speech information for this problem, and our experimental results show that multi-modal approaches significantly outperform the best mono and bimodal models in this task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cutting Edge: Soft Correspondences in Multimodal Scene Parsing

2015-12-01 · ICCV 2015 12 · Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann, Lars Petersson

Exploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the single modality scenario. Existing methods, however, assume that corresponding regions in two modalities have the same…

Scene Parsing

Soft Correspondences in Multimodal Scene Parsing

2017-09-28 · Sarah Taghavi Namin, Mohammad Najafi, Mathieu Salzmann, Lars Petersson

Exploiting multiple modalities for semantic scene parsing has been shown to improve accuracy over the singlemodality scenario. However multimodal datasets often suffer from problems such as data misalignment and label in…

Scene Parsing

Beyond RGB: Very High Resolution Urban Remote Sensing With Multimodal Deep Networks

2017-11-23 · Nicolas Audebert, Bertrand Le Saux, Sébastien Lefèvre

In this work, we investigate various methods to deal with semantic labeling of very high resolution multi-modal remote sensing data. Especially, we study how deep fully convolutional networks can be adapted to deal with …

Semantic Segmentation

Learning from the Best: Active Learning for Wireless Communications

2024-01-23 · Nasim Soltani, Jifan Zhang, Batool Salehi, Debashri Roy 외

Collecting an over-the-air wireless communications training dataset for deep learning-based communication tasks is relatively simple. However, labeling the dataset requires expert involvement and domain knowledge, may in…

Active LearningDeep Learning

MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling

2025-08-24 · Hyeyeon Kim, Sungwoo Han, Jingun Kwon, Hidetaka Kamigaito 외 arxiv

In this study, we introduce a novel cover image generation task that produces both a concise summary and a visually corresponding image from a given text-only document. Because no existing datasets are available for this…

Image Generation