paper-with-me

홈 › Papers

Leveraging OpenFlamingo for Multimodal Embedding Analysis of C2C Car Parts Data

2025-03-20 · Maisha Binte Rashid, Pablo Rivas

In this paper, we aim to investigate the capabilities of multimodal machine learning models, particularly the OpenFlamingo model, in processing a large-scale dataset of consumer-to-consumer (C2C) online posts related to car parts. We have collected data from two platforms, OfferUp and Craigslist, resulting in a dataset of over 1.2 million posts with their corresponding images. The OpenFlamingo model was used to extract embeddings for the text and image of each post. We used $k$-means clustering on the joint embeddings to identify underlying patterns and commonalities among the posts. We have found that most clusters contain a pattern, but some clusters showed no internal patterns. The results provide insight into the fact that OpenFlamingo can be used for finding patterns in large datasets but needs some modification in the architecture according to the dataset.

📄 PDF Abstract BibTeX arXiv:2503.17408

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Utility of Multimodal Large Language Models in Analyzing Chest X-ray with Incomplete Contextual Information

2024-09-20 · Choonghan Kim, Seonhee Cho, Joo Heung Yoon

Background: Large language models (LLMs) are gaining use in clinical settings, but their performance can suffer with incomplete radiology reports. We tested whether multimodal LLMs (using text and images) could improve a…

Large Language Model

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

2024-01-01 · Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin, JianFeng Wang 외

In the evolution of Vision-Language Pre-training, shifting from short-text comprehension to encompassing extended textual contexts is pivotal. Recent autoregressive vision-language models like \cite{flamingo, palme}, lev…

Language ModellingReading ComprehensionText Generation

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

2023-08-02 · Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel 외

We introduce OpenFlamingo, a family of autoregressive vision-language models ranging from 3B to 9B parameters. OpenFlamingo is an ongoing effort to produce an open-source replication of DeepMind's Flamingo models. On sev…

Visual Question AnsweringVisual Question Answering (VQA)

What Makes Multimodal In-Context Learning Work?

2024-04-24 · Folco Bertini Baldassini, Mustafa Shukor, Matthieu Cord, Laure Soulier 외

Large Language Models have demonstrated remarkable performance across various tasks, exhibiting the capacity to swiftly acquire new skills, such as through In-Context Learning (ICL) with minimal demonstration examples. I…

In-Context Learning

Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

2024-02-19 · Christian Schlarmann, Naman Deep Singh, Francesco Croce, Matthias Hein

Multi-modal foundation models like OpenFlamingo, LLaVA, and GPT-4 are increasingly used for various real-world tasks. Prior work has shown that these models are highly vulnerable to adversarial attacks on the vision moda…

Adversarial DefenseMultimodal Deep Learningzero-shot-classificationZero-Shot Learning