A multimodal deep learning framework for scalable content based visual media retrieval
We propose a novel, efficient, modular and scalable framework for content based visual media retrieval systems by leveraging the power of Deep Learning which is flexible to work both for images and videos conjointly and we also introduce an efficient comparison and filtering metric for retrieval. We put forward our findings from critical performance tests comparing our method to the predominant conventional approach to demonstrate the feasibility and efficiency of the proposed solution with best practices, possible improvements that may further augment the ability of retrieval architectures.
Code (2)
Tasks
Multimodal Deep LearningRetrievalSimilar Papers 제목 키워드 기반
Structured Graph Representations for Visual Narrative Reasoning: A Hierarchical Framework for Comics
This paper presents a hierarchical knowledge graph framework for the structured understanding of visual narratives, focusing on multimodal media such as comics. The proposed method decomposes narrative content into multi…
Knowledge GraphsMultimodal ReasoningTowards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models
The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderato…
Few-Shot LearningIn-Context LearningMTikGuard System: A Transformer-Based Multimodal System for Child-Safe Content Moderation on TikTok
With the rapid rise of short-form videos, TikTok has become one of the most influential platforms among children and teenagers, but also a source of harmful content that can affect their perception and behavior. Such con…
Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
Broadcast and media organizations increasingly rely on artificial intelligence to automate the labor-intensive processes of content indexing, tagging, and metadata generation. However, existing AI systems typically opera…
Representation LearningCross-Modal RetrievalAnalyzing Sustainability Messaging in Large-Scale Corporate Social Media
In this work, we introduce a multimodal analysis pipeline that leverages large foundation models in vision and language to analyze corporate social media content, with a focus on sustainability-related communication. Add…