Multimodality in Meta-Learning: A Comprehensive Survey
Meta-learning has gained wide popularity as a training framework that is more data-efficient than traditional machine learning methods. However, its generalization ability in complex task distributions, such as multimodal tasks, has not been thoroughly studied. Recently, some studies on multimodality-based meta-learning have emerged. This survey provides a comprehensive overview of the multimodality-based meta-learning landscape in terms of the methodologies and applications. We first formalize the definition of meta-learning in multimodality, along with the research challenges in this growing field, such as how to enrich the input in few-shot learning (FSL) or zero-shot learning (ZSL) in multimodal scenarios and how to generalize the models to new tasks. We then propose a new taxonomy to discuss typical meta-learning algorithms in multimodal tasks systematically. We investigate the contributions of related papers and summarize them by our taxonomy. Finally, we propose potential research directions for this promising field.
Code (0)
등록된 구현이 없습니다.
Tasks
Few-Shot LearningMeta-LearningSurveyZero-Shot LearningSimilar Papers 제목 키워드 기반
Leveraging Multimodality for Real-Time Classification of Transients and Variables found by the Zwicky Transient Facility
Modern time-domain surveys such as the Zwicky Transient Facility (ZTF) generate hundreds of thousands of alerts each night, making real-time decisions for follow-up observations a central challenge in time-domain astrono…
Multimodality Representation Learning: A Survey on Evolution, Pretraining and Its Applications
Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Questi…
Question AnsweringRepresentation LearningRetrievalSurvey+3Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality
Remote sensing (RS) techniques are increasingly crucial for deepening our understanding of the planet. As the volume and diversity of RS data continue to grow exponentially, there is an urgent need for advanced data mode…
Text to Image Generation and Editing: A Survey
Text-to-image generation (T2I) refers to the text-guided generation of high-quality images. In the past few years, T2I has attracted widespread attention and numerous works have emerged. In this survey, we comprehensivel…
Image GenerationMambaSurveytext-guided-generation+2Video Quality Assessment: A Comprehensive Survey
Video quality assessment (VQA) is an important processing task, aiming at predicting the quality of videos in a manner highly consistent with human judgments of perceived quality. Traditional VQA models based on natural …
BenchmarkingSurveyVideo Quality AssessmentVisual Question Answering (VQA)