paper-with-me

홈 › Papers

Using Multiple Input Modalities Can Improve Data-Efficiency and O.O.D. Generalization for ML with Satellite Imagery

2025-07-15 · Arjun Rao, Esther Rolf arxiv

A large variety of geospatial data layers is available around the world ranging from remotely-sensed raster data like satellite imagery, digital elevation models, predicted land cover maps, and human-annotated data, to data derived from environmental sensors such as air temperature or wind speed data. A large majority of machine learning models trained on satellite imagery (SatML), however, are designed primarily for optical input modalities such as multi-spectral satellite imagery. To better understand the value of using other input modalities alongside optical imagery in supervised learning settings, we generate augmented versions of SatML benchmark tasks by appending additional geographic data layers to datasets spanning classification, regression, and segmentation. Using these augmented datasets, we find that fusing additional geographic inputs with optical imagery can significantly improve SatML model performance. Benefits are largest in settings where labeled data are limited and in geographic out-of-sample settings, suggesting that multi-modal inputs may be especially valuable for data-efficiency and out-of-sample performance of SatML models. Surprisingly, we find that hard-coded fusion strategies outperform learned variants, with interesting implications for future work.

📄 PDF Abstract BibTeX arXiv:2507.13385

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

M3ER: Multiplicative Multimodal Emotion Recognition Using Facial, Textual, and Speech Cues

2019-11-09 · Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera 외

We present M3ER, a learning-based method for emotion recognition from multiple input modalities. Our approach combines cues from multiple co-occurring modalities (such as face, text, and speech) and also is more robust t…

Emotion RecognitionMultimodal Emotion Recognition

FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation

2024-05-08 · Xuehai He, Jian Zheng, Jacob Zhiyuan Fang, Robinson Piramuthu 외

Controllable text-to-image (T2I) diffusion models generate images conditioned on both text prompts and semantic inputs of other modalities like edge maps. Nevertheless, current controllable T2I methods commonly face chal…

Image GenerationText to Image GenerationText-to-Image Generation

Multimodal Object Detection by Channel Switching and Spatial Attention

2023-06-18 · Conference on Computer Vision and Pattern Recognition (CVPR) 2023 6 · Yue Cao, Junchi Bin, Jozsef Hamari, Erik Blasch 외

Multimodal object detection has attracted great attention in recent years since the information specific to different modalities can complement each other and effectively improve the accuracy and stability of the detecti…

Multispectral Object Detectionobject-detectionObject DetectionPedestrian Detection

You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language Model

2022-11-21 · CVPR 2023 1 · Shengkun Tang, Yaqing Wang, Zhenglun Kong, Tianchi Zhang 외

Large-scale Transformer models bring significant improvements for various downstream vision language tasks with a unified architecture. The performance improvements come with increasing model size, resulting in slow infe…

DecoderLanguage ModelingLanguage Modelling

Modality-Agnostic Learning for Medical Image Segmentation Using Multi-modality Self-distillation

2023-06-06 · Qisheng He, Nicholas Summerfield, Ming Dong, Carri Glide-Hurst

Medical image segmentation of tumors and organs at risk is a time-consuming yet critical process in the clinic that utilizes multi-modality imaging (e.g, different acquisitions, data types, and sequences) to increase seg…

Image SegmentationMedical Image SegmentationRepresentation LearningSegmentation+1