paper-with-me

홈 › Papers

Can Text-to-image Model Assist Multi-modal Learning for Visual Recognition with Visual Modality Missing?

2024-02-14 · Tiantian Feng, Daniel Yang, Digbalay Bose, Shrikanth Narayanan

Multi-modal learning has emerged as an increasingly promising avenue in vision recognition, driving innovations across diverse domains ranging from media and education to healthcare and transportation. Despite its success, the robustness of multi-modal learning for visual recognition is often challenged by the unavailability of a subset of modalities, especially the visual modality. Conventional approaches to mitigate missing modalities in multi-modal learning rely heavily on algorithms and modality fusion schemes. In contrast, this paper explores the use of text-to-image models to assist multi-modal learning. Specifically, we propose a simple but effective multi-modal learning framework GTI-MM to enhance the data efficiency and model robustness against missing visual modality by imputing the missing data with generative transformers. Using multiple multi-modal datasets with visual recognition tasks, we present a comprehensive analysis of diverse conditions involving missing visual modality in data, including model training. Our findings reveal that synthetic images benefit training data efficiency with visual data missing in training and improve model robustness with visual data missing involving training and testing. Moreover, we demonstrate GTI-MM is effective with lower generation quantity and simple prompt techniques.

📄 PDF Abstract BibTeX arXiv:2402.09036

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MMIU: Dataset for Visual Intent Understanding in Multimodal Assistants

2021-10-13 · Alkesh Patel, Joel Ruben Antony Moniz, Roman Nguyen, Nick Tzou 외

In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. Current digital assistants take spoken in…

intent-classificationIntent ClassificationQuestion AnsweringQuestion Generation+3

Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant

2024-10-24 · Abhirama Subramanyam Penamakuri, Anand Mishra

We revisit knowledge-aware text-based visual question answering, also known as Text-KVQA, in the light of modern advancements in large multimodal models (LMMs), and make the following contributions: (i) We propose VisTEL…

Entity LinkingQuestion AnsweringVisual Question Answering

Multi-modal dialog for browsing large visual catalogs using exploration-exploitation paradigm in a joint embedding space

2019-01-28 · Indrani Bhattacharya, Arkabandhu Chowdhury, Vikas Raykar

We present a multi-modal dialog system to assist online shoppers in visually browsing through large catalogs. Visual browsing is different from visual search in that it allows the user to explore the wide range of produc…

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

2026-05-12 · Zhuoyu Cai, Dou Quan, Ning Huyan, Pei He 외 arxiv

Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However, optical-SAR registration remains challenging under large geometric d…

Image Registration

Generative Visual Instruction Tuning

2024-06-17 · Jefferson Hernandez, Ruben Villegas, Vicente Ordonez

We propose to use automatically generated instruction-following data to improve the zero-shot capabilities of a large multimodal model with additional support for generative and image editing tasks. We achieve this by cu…

Image GenerationImage-text matchingInstruction FollowingLanguage Modeling+4