paper-with-me

Papers

Efficient Large-Scale Multi-Modal Classification

2018-02-06 · D. Kiela, E. Grave, A. Joulin, T. Mikolov

While the incipient internet was largely text-based, the modern digital world is becoming increasingly multi-modal. Here, we examine multi-modal classification where one modality is discrete, e.g. text, and the other is continuous, e.g. visual representations transferred from a convolutional neural network. In particular, we focus on scenarios where we have to be able to classify large quantities of data quickly. We investigate various methods for performing multi-modal fusion and analyze their trade-offs in terms of classification accuracy and computational efficiency. Our findings indicate that the inclusion of continuous information improves performance over text-only on a range of multi-modal classification tasks, even with simple fusion methods. In addition, we experiment with discretizing the continuous features in order to speed up and simplify the fusion process even further. Our results show that fusion with discretized features outperforms text-only classification, at a fraction of the computational cost of full multi-modal fusion, with the additional benefit of improved interpretability.

📄 PDF Abstract BibTeX arXiv:1802.02892

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationComputational EfficiencyGeneral ClassificationMulti-modal Classification

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Towards Good Practices for Multi-modal Fusion in Large-scale Video Classification

2018-09-16 · Jinlai Liu, Zehuan Yuan, Changhu Wang

Leveraging both visual frames and audio has been experimentally proven effective to improve large-scale video classification. Previous research on video classification mainly focuses on the analysis of visual content amo…

ClassificationGeneral ClassificationVideo Classification

X-ModalNet: A Semi-Supervised Deep Cross-Modal Network for Classification of Remote Sensing Data

2020-06-24 · Danfeng Hong, Naoto Yokoya, Gui-Song Xia, Jocelyn Chanussot 외

This paper addresses the problem of semi-supervised transfer learning with limited cross-modality data in remote sensing. A large amount of multi-modal earth observation images, such as multispectral imagery (MSI) or syn…

Earth ObservationGeneral ClassificationTransfer Learning

Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching

2025-12-03 · Wei Chee Yew, Hailun Xu, Sanjay Saha, Xiaotian Fan 외 arxiv

Content moderation remains a critical yet challenging task for large-scale user-generated video platforms, especially in livestreaming environments where moderation must be timely, multimodal, and robust to evolving form…

Multi Task Learning based Framework for Multimodal Classification

2021-06-01 · NAACL (maiworkshop) 2021 6 · Danting Zeng

Large-scale multi-modal classification aim to distinguish between different multi-modal data, and it has drawn dramatically attentions since last decade. In this paper, we propose a multi-task learning-based framework fo…

ClassificationMulti-modal ClassificationMulti-Task Learning

Image and Encoded Text Fusion for Multi-Modal Classification

2018-10-03 · Ignazio Gallo, Alessandro Calefati, Shah Nawaz, Muhammad Kamran Janjua

Multi-modal approaches employ data from multiple input streams such as textual and visual domains. Deep neural networks have been successfully employed for these approaches. In this paper, we present a novel multi-modal …

ClassificationGeneral ClassificationMulti-modal Classification