Papers Image-text Classification
“Image-text Classification” 태그가 달린 논문 13편 · 필터 해제
Unified Generative and Discriminative Training for Multi-modal Large Language Models
In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle various complex tasks, yet issues such …
Dynamic Time WarpingImage-text ClassificationLanguage ModelingLanguage Modelling+4Multimodal Quantum Natural Language Processing: A Novel Framework for using Quantum Methods to Analyse Real Data
Despite significant advances in quantum computing across various domains, research on applying quantum approaches to language compositionality - such as modeling linguistic structures and interactions - remains limited. …
Data IntegrationImage-text ClassificationLanguage ModelingLanguage Modelling+2Leveraging Foundation Models for Multi-modal Federated Learning with Incomplete Modality
Federated learning (FL) has obtained tremendous progress in providing collaborative training solutions for distributed data silos with privacy guarantees. However, few existing works explore a more realistic scenario whe…
Federated LearningImage-text ClassificationModality completiontext-classification+2Robust Latent Representation Tuning for Image-text Classification
Large models have demonstrated exceptional generalization capabilities in computer vision and natural language processing. Recent efforts have focused on enhancing these models with multimodal processing abilities. Howev…
ClassificationImage-text Classificationtext-classificationText ClassificationContinuous Geometry-Aware Graph Diffusion via Hyperbolic Neural PDE
While Hyperbolic Graph Neural Network (HGNN) has recently emerged as a powerful tool dealing with hierarchical graph data, the limitations of scalability and efficiency hinder itself from generalizing to deep models. In …
Graph Neural NetworkImage-text ClassificationLink PredictionNode Classification+2GIST: Generating Image-Specific Text for Fine-grained Object Classification
Recent vision-language models outperform vision-only models on many image classification tasks. However, because of the absence of paired text/image descriptions, it remains difficult to fine-tune these models for fine-g…
ClassificationFine-Grained Image Classificationimage-classificationImage Classification+6UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning
Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the int…
Contrastive LearningImage-text Classificationtext-classificationText ClassificationTowards Unifying Medical Vision-and-Language Pre-training via Soft Prompts
Medical vision-and-language pre-training (Med-VLP) has shown promising improvements on many downstream medical tasks owing to its applicability to extracting generic representations from medical images and texts. Practic…
Image RetrievalImage-text ClassificationImage to textQuestion Answering+7DIFFormer: Scalable (Graph) Transformers Induced by Energy Constrained Diffusion
Real-world data generation often involves complex inter-dependencies among instances, violating the IID-data hypothesis of standard learning paradigms and posing a challenge for uncovering the geometric structures for le…
Image-text ClassificationNode Classificationtext-classificationText ClassificationGLAMI-1M: A Multilingual Image-Text Fashion Dataset
We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 …
ClassificationImage GenerationImage-text ClassificationMultilingual Image-Text Classification+2Context-Aware Compilation of DNN Training Pipelines across Edge and Cloud
Empowered by machine learning, edge devices including smartphones, wearable, and IoT devices have become growingly intelligent, raising conflicts with the limited resource. On-device model personalization is particularly…
Feature CompressionImage ClassificationImage GenerationImage-text Classification+3CMA-CLIP: Cross-Modality Attention CLIP for Image-Text Classification
Modern Web systems such as social media and e-commerce contain rich contents expressed in images and text. Leveraging information from multi-modalities can improve the performance of machine learning tasks such as classi…
AttributeImage-text ClassificationMultimodal Text and Image Classificationtext-classification+1Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!
Modeling expressive cross-modal interactions seems crucial in multimodal tasks, such as visual question answering. However, sometimes high-performing black-box algorithms turn out to be mostly exploiting unimodal signals…
DiagnosticImage-text ClassificationQuestion Answeringtext-classification+3