paper-with-me

Papers

How to find a good image-text embedding for remote sensing visual question answering?

2021-09-24 · Christel Chappuis, Sylvain Lobry, Benjamin Kellenberger, Bertrand Le Saux, Devis Tuia

Visual question answering (VQA) has recently been introduced to remote sensing to make information extraction from overhead imagery more accessible to everyone. VQA considers a question (in natural language, therefore easy to formulate) about an image and aims at providing an answer through a model based on computer vision and natural language processing methods. As such, a VQA model needs to jointly consider visual and textual features, which is frequently done through a fusion step. In this work, we study three different fusion methodologies in the context of VQA for remote sensing and analyse the gains in accuracy with respect to the model complexity. Our findings indicate that more complex fusion mechanisms yield an improved performance, yet that seeking a trade-of between model complexity and performance is worthwhile in practice.

📄 PDF Abstract BibTeX arXiv:2109.11848

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Txt2Img-MHN: Remote Sensing Image Generation from Text Using Modern Hopfield Networks

2022-08-08 · Yonghao Xu, Weikang Yu, Pedram Ghamisi, Michael Kopp 외

The synthesis of high-resolution remote sensing images based on text descriptions has great potential in many practical application scenarios. Although deep neural networks have achieved great success in many important r…

Image GenerationText to Image GenerationText-to-Image Generationzero-shot-classification+1

Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion

2026-07-28 · Junhyuk Heo, Junghwan Park arxiv

Open-vocabulary segmentation labels arbitrary categories from a text query without per-class training, yet on remote sensing imagery it underperforms on categories it handles reliably elsewhere. We find that much of this…

Effective Use of Dilated Convolutions for Segmenting Small Object Instances in Remote Sensing Imagery

2017-09-01 · Ryuhei Hamaguchi, Aito Fujita, Keisuke Nemoto, Tomoyuki Imaizumi 외

Thanks to recent advances in CNNs, solid improvements have been made in semantic segmentation of high resolution remote sensing imagery. However, most of the previous works have not fully taken into account the specific …

Semantic Segmentation

Training general representations for remote sensing using in-domain knowledge

2020-09-30 · Maxim Neumann, André Susano Pinto, Xiaohua Zhai, Neil Houlsby

Automatically finding good and general remote sensing representations allows to perform transfer learning on a wide range of applications - improving the accuracy and reducing the required number of training samples. Thi…

Representation LearningTransfer Learning

MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion Models

2024-12-19 · Jing Zhao, Heliang Zheng, Chaoyue Wang, Long Lan 외

Large-scale text-to-image diffusion models, (e.g., DALL-E, SDXL) are capable of generating famous persons by simply referring to their names. Is it possible to make such models generate generic identities as simple as th…