Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries
The complexity of the visual world creates significant challenges for comprehensive visual understanding. In spite of recent successes in visual recognition, today's vision systems would still struggle to deal with visual queries that require a deeper reasoning. We propose a knowledge base (KB) framework to handle an assortment of visual queries, without the need to train new classifiers for new tasks. Building such a large-scale multimodal KB presents a major challenge of scalability. We cast a large-scale MRF into a KB representation, incorporating visual, textual and structured data, as well as their diverse relations. We introduce a scalable knowledge base construction system that is capable of building a KB with half billion variables and millions of parameters in a few hours. Our system achieves competitive results compared to purpose-built models on standard recognition and retrieval tasks, while exhibiting greater flexibility in answering richer visual queries.
Code (0)
등록된 구현이 없습니다.
Tasks
Knowledge Base ConstructionRetrievalSimilar Papers 제목 키워드 기반
Construction and Applications of Billion-Scale Pre-Trained Multimodal Business Knowledge Graph
Business Knowledge Graphs (KGs) are important to many enterprises today, providing factual knowledge and structured data that steer many products and make them more intelligent. Despite their promising benefits, building…
Knowledge GraphsBe My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
Large Language Models (LLMs) have demonstrated remarkable capabilities in challenging, knowledge-intensive reasoning tasks. However, extending LLMs to perceive and reason over a new modality (e.g., vision), often require…
Multimodal ReasoningEmpowering Time Series Analysis with Large-Scale Multimodal Pretraining
While existing time series foundation models primarily rely on large-scale unimodal pretraining, they lack complementary modalities to enhance time series understanding. Building multimodal foundation models is a natural…
Time Series ForecastingTime Series AnalysisAnomaly DetectionConnecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
Few-shot learning (FSL) addresses the challenge of classifying novel classes with limited training samples. While some methods leverage semantic knowledge from smaller-scale models to mitigate data scarcity, these approa…
Few-Shot LearningWhen Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
Vision-language models (VLMs) have achieved remarkable success across multimodal tasks, yet their substantial computational demands hinder efficient deployment. Knowledge distillation (KD) has emerged as a powerful appro…
Visual Question AnsweringKnowledge Distillation