Deep Multimodal Transfer-Learned Regression in Data-Poor Domains
In many real-world applications of deep learning, estimation of a target may rely on various types of input data modes, such as audio-video, image-text, etc. This task can be further complicated by a lack of sufficient data. Here we propose a Deep Multimodal Transfer-Learned Regressor (DMTL-R) for multimodal learning of image and feature data in a deep regression architecture effective at predicting target parameters in data-poor domains. Our model is capable of fine-tuning a given set of pre-trained CNN weights on a small amount of training image data, while simultaneously conditioning on feature information from a complimentary data mode during network training, yielding more accurate single-target or multi-target regression than can be achieved using the images or the features alone. We present results using phase-field simulation microstructure images with an accompanying set of physical features, using pre-trained weights from various well-known CNN architectures, which demonstrate the efficacy of the proposed multimodal approach.
Code (1)
Tasks
Multi-target regressionregressionSimilar Papers 제목 키워드 기반
VideoAdviser: Video Knowledge Distillation for Multimodal Transfer Learning
Multimodal transfer learning aims to transform pretrained representations of diverse modalities into a common domain space for effective multimodal fusion. However, conventional systems are typically built on the assumpt…
Knowledge DistillationregressionSentiment AnalysisTransfer LearningWhy Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks
In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations. Despite its success in large language models, the extension of ICL to multimodal settings remains poorly understood in ter…
The Common Intuition to Transfer Learning Can Win or Lose: Case Studies for Linear Regression
We study a fundamental transfer learning process from source to target linear regression tasks, including overparameterized settings where there are more learned parameters than data samples. The target task learning is …
PhilosophyregressionTransfer LearningBounded-Compute Multimodal Regression for Product-Rating Prediction
Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generation and dynamic visual processing is poorly matched to scalar regressi…
Text GenerationVersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment
Large Multimodal Models (LMMs) have shown promise for video quality assessment, but most methods still predict an absolute score for each video. Such pointwise supervision often mixes perceptual quality with dataset-spec…
Video Quality AssessmentDomain GeneralizationRelational Reasoning