paper-with-me

Papers

A Unified Continuous Learning Framework for Multi-modal Knowledge Discovery and Pre-training

2022-06-11 · Zhihao Fan, Zhongyu Wei, Jingjing Chen, Siyuan Wang, Zejun Li, Jiarong Xu, Xuanjing Huang

Multi-modal pre-training and knowledge discovery are two important research topics in multi-modal machine learning. Nevertheless, none of existing works make attempts to link knowledge discovery with knowledge guided multi-modal pre-training. In this paper, we propose to unify them into a continuous learning framework for mutual improvement. Taking the open-domain uni-modal datasets of images and texts as input, we maintain a knowledge graph as the foundation to support these two tasks. For knowledge discovery, a pre-trained model is used to identify cross-modal links on the graph. For model pre-training, the knowledge graph is used as the external knowledge to guide the model updating. These two steps are iteratively performed in our framework for continuous learning. The experimental results on MS-COCO and Flickr30K with respect to both knowledge discovery and the pre-trained model validate the effectiveness of our framework.

📄 PDF Abstract BibTeX arXiv:2206.05555

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Harmony: A Unified Framework for Modality Incremental Learning

2025-04-17 · Yaguang Song, Xiaoshan Yang, Dongmei Jiang, YaoWei Wang 외

Incremental learning aims to enable models to continuously acquire knowledge from evolving data streams while preserving previously learned capabilities. While current research predominantly focuses on unimodal increment…

Incremental Learning

Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing

2026-05-13 · Blaise Delattre, Hengyu Wu, Paul Caillon, Wei Yang Bryan Lim 외 arxiv

Randomized smoothing provides strong, model-agnostic robustness certificates, but existing guarantees are limited to single modalities, treating continuous and discrete inputs in isolation. This limitation becomes critic…

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding

2026-04-15 · Yibo Jiang, Tao Wu, Rui Jiang, Yehao Lu 외 arxiv

Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability significa…

Visual Reasoning

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

2026-04-09 · Luozheng Qin, Jia Gong, Qian Qiao, Tianjiao Li 외 arxiv

Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher computational costs than understanding, particularly for video. This i…

multimodal generationVideo GenerationText Generation

VL-KGE: Vision-Language Models Meet Knowledge Graph Embeddings

2026-03-02 · Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg 외 arxiv

Real-world multimodal knowledge graphs (MKGs) are inherently heterogeneous, modeling entities that are associated with diverse modalities. Traditional knowledge graph embedding (KGE) methods excel at learning continuous …

Knowledge Graph EmbeddingKnowledge GraphsLink Prediction