paper-with-me

Papers Image-text Classification

“Image-text Classification” 태그가 달린 논문 13편 · 필터 해제

Unified Generative and Discriminative Training for Multi-modal Large Language Models

2024-11-01 · Wei Chow, Juncheng Li, Qifan Yu, Kaihang Pan 외

In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle various complex tasks, yet issues such …

Dynamic Time WarpingImage-text ClassificationLanguage ModelingLanguage Modelling+4

Multimodal Quantum Natural Language Processing: A Novel Framework for using Quantum Methods to Analyse Real Data

2024-10-29 · Hala Hawashin

Despite significant advances in quantum computing across various domains, research on applying quantum approaches to language compositionality - such as modeling linguistic structures and interactions - remains limited. …

Data IntegrationImage-text ClassificationLanguage ModelingLanguage Modelling+2

Leveraging Foundation Models for Multi-modal Federated Learning with Incomplete Modality

2024-06-16 · Liwei Che, Jiaqi Wang, Xinyue Liu, Fenglong Ma

Federated learning (FL) has obtained tremendous progress in providing collaborative training solutions for distributed data silos with privacy guarantees. However, few existing works explore a more realistic scenario whe…

Federated LearningImage-text ClassificationModality completiontext-classification+2

Robust Latent Representation Tuning for Image-text Classification

2024-06-10 · Hao Sun, Yu Song

Large models have demonstrated exceptional generalization capabilities in computer vision and natural language processing. Recent efforts have focused on enhancing these models with multimodal processing abilities. Howev…

ClassificationImage-text Classificationtext-classificationText Classification

Continuous Geometry-Aware Graph Diffusion via Hyperbolic Neural PDE

2024-06-03 · Jiaxu Liu, Xinping Yi, Sihao Wu, Xiangyu Yin 외

While Hyperbolic Graph Neural Network (HGNN) has recently emerged as a powerful tool dealing with hierarchical graph data, the limitations of scalability and efficiency hinder itself from generalizing to deep models. In …

Graph Neural NetworkImage-text ClassificationLink PredictionNode Classification+2

GIST: Generating Image-Specific Text for Fine-grained Object Classification

2023-07-21 · Kathleen M. Lewis, Emily Mu, Adrian V. Dalca, John Guttag

Recent vision-language models outperform vision-only models on many image classification tasks. However, because of the absence of paired text/image descriptions, it remains difficult to fine-tune these models for fine-g…

ClassificationFine-Grained Image Classificationimage-classificationImage Classification+6

UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning

2023-05-16 · Heqing Zou, Meng Shen, Chen Chen, Yuchen Hu 외

Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based multimodal fusion methods ignore the int…

Contrastive LearningImage-text Classificationtext-classificationText Classification

Towards Unifying Medical Vision-and-Language Pre-training via Soft Prompts

2023-02-17 · ICCV 2023 1 · Zhihong Chen, Shizhe Diao, Benyou Wang, Guanbin Li 외

Medical vision-and-language pre-training (Med-VLP) has shown promising improvements on many downstream medical tasks owing to its applicability to extracting generic representations from medical images and texts. Practic…

Image RetrievalImage-text ClassificationImage to textQuestion Answering+7

DIFFormer: Scalable (Graph) Transformers Induced by Energy Constrained Diffusion

2023-01-23 · Qitian Wu, Chenxiao Yang, Wentao Zhao, Yixuan He 외

Real-world data generation often involves complex inter-dependencies among instances, violating the IID-data hypothesis of standard learning paradigms and posing a challenge for uncovering the geometric structures for le…

Image-text ClassificationNode Classificationtext-classificationText Classification

GLAMI-1M: A Multilingual Image-Text Fashion Dataset

2022-11-17 · BMVC 2022 11 · Vaclav Kosar, Antonín Hoskovec, Milan Šulc, Radek Bartyzal

We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 …

ClassificationImage GenerationImage-text ClassificationMultilingual Image-Text Classification+2

Context-Aware Compilation of DNN Training Pipelines across Edge and Cloud

2021-12-30 · Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2021 12 · Dixi Yao, Liyao Xiang, Zifan Wang, Jiayu Xu 외

Empowered by machine learning, edge devices including smartphones, wearable, and IoT devices have become growingly intelligent, raising conflicts with the limited resource. On-device model personalization is particularly…

Feature CompressionImage ClassificationImage GenerationImage-text Classification+3

CMA-CLIP: Cross-Modality Attention CLIP for Image-Text Classification

2021-12-07 · Huidong Liu, Shaoyuan Xu, Jinmiao Fu, Yang Liu 외

Modern Web systems such as social media and e-commerce contain rich contents expressed in images and text. Leveraging information from multi-modalities can improve the performance of machine learning tasks such as classi…

AttributeImage-text ClassificationMultimodal Text and Image Classificationtext-classification+1

Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!

2020-10-13 · EMNLP 2020 11 · Jack Hessel, Lillian Lee

Modeling expressive cross-modal interactions seems crucial in multimodal tasks, such as visual question answering. However, sometimes high-performing black-box algorithms turn out to be mostly exploiting unimodal signals…

DiagnosticImage-text ClassificationQuestion Answeringtext-classification+3
1–13 / 13