paper-with-me

홈 › Papers

More Diverse Means Better: Multimodal Deep Learning Meets Remote Sensing Imagery Classification

2020-08-12 · Danfeng Hong, Lianru Gao, Naoto Yokoya, Jing Yao, Jocelyn Chanussot, Qian Du, Bing Zhang

Classification and identification of the materials lying over or beneath the Earth's surface have long been a fundamental but challenging research topic in geoscience and remote sensing (RS) and have garnered a growing concern owing to the recent advancements of deep learning techniques. Although deep networks have been successfully applied in single-modality-dominated classification tasks, yet their performance inevitably meets the bottleneck in complex scenes that need to be finely classified, due to the limitation of information diversity. In this work, we provide a baseline solution to the aforementioned difficulty by developing a general multimodal deep learning (MDL) framework. In particular, we also investigate a special case of multi-modality learning (MML) -- cross-modality learning (CML) that exists widely in RS image classification applications. By focusing on "what", "where", and "how" to fuse, we show different fusion strategies as well as how to train deep networks and build the network architecture. Specifically, five fusion architectures are introduced and developed, further being unified in our MDL framework. More significantly, our framework is not only limited to pixel-wise classification tasks but also applicable to spatial information modeling with convolutional neural networks (CNNs). To validate the effectiveness and superiority of the MDL framework, extensive experiments related to the settings of MML and CML are conducted on two different multimodal RS datasets. Furthermore, the codes and datasets will be available at https://github.com/danfenghong/IEEE_TGRS_MDL-RS, contributing to the RS community.

📄 PDF Abstract BibTeX arXiv:2008.05457

Code (1)

danfenghong/IEEE_TGRS_MDL-RS 공식 구현 tf

Tasks

ClassificationGeneral Classificationimage-classificationImage ClassificationMultimodal Deep Learning

Methods 이 논문이 사용한 방법론

MDL Minimum Description Length provides a criterion for the selection of models, regardless of their complexity, without the restrictive assumption that the data form a sample…

Similar Papers 제목 키워드 기반

MixMAS: A Framework for Sampling-Based Mixer Architecture Search for Multimodal Fusion and Learning

2024-12-24 · Abdelmadjid Chergui, Grigor Bezirganyan, Sana Sellami, Laure Berti-ÉQuille 외

Choosing a suitable deep learning architecture for multimodal data fusion is a challenging task, as it requires the effective integration and processing of diverse data types, each with distinct structures and characteri…

Benchmarking

EmoLLM: Multimodal Emotional Understanding Meets Large Language Models

2024-06-24 · Qu Yang, Mang Ye, Bo Du

Multi-modal large language models (MLLMs) have achieved remarkable performance on objective multimodal perception tasks, but their ability to interpret subjective, emotionally nuanced multimodal content remains largely u…

Emotional Intelligence

Fusion Meets Diverse Conditions: A High-diversity Benchmark and Baseline for UAV-based Multimodal Object Detection with Condition Cues

2025-10-15 · Chen Chen, Kangcheng Bin, Ting Hu, Jiahao Qi 외 arxiv

Unmanned aerial vehicles (UAV)-based object detection with visible (RGB) and infrared (IR) images facilitates robust around-the-clock detection, driven by advancements in deep learning techniques and the availability of …

Object Detection

Alquist 5.0: Dialogue Trees Meet Generative Models. A Novel Approach for Enhancing SocialBot Conversations

2023-10-24 · Ondřej Kobza, Jan Čuhel, Tommaso Gargiani, David Herel 외

We present our SocialBot -- Alquist~5.0 -- developed for the Alexa Prize SocialBot Grand Challenge~5. Building upon previous versions of our system, we introduce the NRG Barista and outline several innovative approaches …

ChatGPT Meets Iris Biometrics

2024-08-09 · Parisa Farmanifard, Arun Ross

This study utilizes the advanced capabilities of the GPT-4 multimodal Large Language Model (LLM) to explore its potential in iris recognition - a field less common and more specialized than face recognition. By focusing …

Face RecognitionIris RecognitionLanguage ModelingLanguage Modelling+3