paper-with-me

홈 › Papers

Efficient Generalization via Multimodal Co-Training under Data Scarcity and Distribution Shift

2025-10-08 · Tianyu Bell Pan, Damon L. Woodard arxiv

This paper explores a multimodal co-training framework designed to enhance model generalization in situations where labeled data is limited and distribution shifts occur. We thoroughly examine the theoretical foundations of this framework, deriving conditions under which the use of unlabeled data and the promotion of agreement between classifiers for different modalities lead to significant improvements in generalization. We also present a convergence analysis that confirms the effectiveness of iterative co-training in reducing classification errors. In addition, we establish a novel generalization bound that, for the first time in a multimodal co-training context, decomposes and quantifies the distinct advantages gained from leveraging unlabeled multimodal data, promoting inter-view agreement, and maintaining conditional view independence. Our findings highlight the practical benefits of multimodal co-training as a structured approach to developing data-efficient and robust AI systems that can effectively generalize in dynamic, real-world environments. The theoretical foundations are examined in dialogue with, and in advance of, established co-training principles.

📄 PDF Abstract BibTeX arXiv:2510.07509

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Breaking the Data Barrier -- Building GUI Agents Through Task Generalization

2025-04-14 · Junlei Zhang, Zichen Ding, Chang Ma, Zijie Chen 외

Graphical User Interface (GUI) agents offer cross-platform solutions for automating complex digital tasks, with significant potential to transform productivity workflows. However, their performance is often constrained b…

Mathematical ReasoningMultimodal ReasoningTransfer Learning

A Shared Encoder Approach to Multimodal Representation Learning

2025-03-03 · Shuvendu Roy, Franklin Ogidi, Ali Etemad, Elham Dolatabadi 외

Multimodal representation learning has demonstrated remarkable potential in enabling models to process and integrate diverse data modalities, such as text and images, for improved understanding and performance. While the…

Representation Learning

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis

2026-01-11 · Jiazhang Liang, Jianheng Dai, Miaosen Luo, Menghua Jiang 외 arxiv

Multimodal large language models have demonstrated strong ability in capturing semantic representations for multimodal sentiment analysis. Their capacity to learn stable and generalizable multimodal features is limited, …

Multimodal Sentiment AnalysisData Augmentation

M$^3$Searcher: Modular Multimodal Information Seeking Agency with Retrieval-Oriented Reasoning

2026-01-14 · Xiaohan Yu, Chao Feng, Lang Mei, Chong Chen arxiv

Recent advances in DeepResearch-style agents have demonstrated strong capabilities in autonomous information acquisition and synthesize from real-world web environments. However, existing approaches remain fundamentally …

Contrastive Learning of English Language and Crystal Graphs for Multimodal Representation of Materials Knowledge

2025-02-23 · Yang Jeong Park, Mayank Kumaran, Chia-Wei Hsu, Elsa Olivetti 외

Artificial intelligence (AI) is increasingly used for the inverse design of materials, such as crystals and molecules. Existing AI research on molecules has integrated chemical structures of molecules with textual knowle…

Contrastive LearningZero-shot Generalization