paper-with-me

홈 › Papers

xGen-MM (BLIP-3): A Family of Open Large Multimodal Models

2024-08-16 · Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S Ryoo, Shrikant Kendre, Jieyu Zhang, Can Qin, Shu Zhang, Chia-Chih Chen, Ning Yu, Juntao Tan, Tulika Manoj Awalgaonkar, Shelby Heinecke, Huan Wang, Yejin Choi, Ludwig Schmidt, Zeyuan Chen, Silvio Savarese, Juan Carlos Niebles, Caiming Xiong, ran Xu

This report introduces xGen-MM (also known as BLIP-3), a framework for developing Large Multimodal Models (LMMs). The framework comprises meticulously curated datasets, a training recipe, model architectures, and a resulting suite of LMMs. xGen-MM, short for xGen-MultiModal, expands the Salesforce xGen initiative on foundation AI models. Our models undergo rigorous evaluation across a range of tasks, including both single and multi-image benchmarks. Our pre-trained base model exhibits strong in-context learning capabilities and the instruction-tuned model demonstrates competitive performance among open-source LMMs with similar model sizes. In addition, we introduce a safety-tuned model with DPO, aiming to mitigate harmful behaviors such as hallucinations and improve safety. We open-source our models, curated large-scale datasets, and our fine-tuning codebase to facilitate further advancements in LMM research. Associated resources will be available on our project page above.

📄 PDF Abstract BibTeX arXiv:2408.08872

Code (1)

zzxslp/som-llava pytorch

Tasks

In-Context Learning

Methods 이 논문이 사용한 방법론

DPO 설명 없음
BASE 설명 없음

Similar Papers 제목 키워드 기반

xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

2024-10-21 · Michael S. Ryoo, Honglu Zhou, Shrikant Kendre, Can Qin 외

We present xGen-MM-Vid (BLIP-3-Video): a multimodal language model for videos, particularly designed to efficiently capture temporal information over multiple frames. BLIP-3-Video takes advantage of the 'temporal encoder…

Language ModelingLanguage ModellingQuestion AnsweringVideo Question Answering

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

2025-05-14 · Jiuhai Chen, Zhiyang Xu, Xichen Pan, Yushi Hu 외

Unifying image understanding and generation has gained growing attention in recent research on multimodal models. Although design choices for image understanding have been extensively studied, the optimal model architect…

Image Generation

xGen-small Technical Report

2025-05-10 · Erik Nijkamp, Bo Pang, Egor Pakhomov, Akash Gokul 외

We introduce xGen-small, a family of 4B and 9B Transformer decoder models optimized for long-context applications. Our vertically integrated pipeline unites domain-balanced, frequency-aware data curation; multi-stage pre…

DecoderMath

XGen-7B Technical Report

2023-09-07 · Erik Nijkamp, Tian Xie, Hiroaki Hayashi, Bo Pang 외

Large Language Models (LLMs) have become ubiquitous across various domains, transforming the way we interact with information and conduct research. However, most high-performing LLMs remain confined behind proprietary wa…

2k8k

TT-BLIP: Enhancing Fake News Detection Using BLIP and Tri-Transformer

2024-03-19 · Eunjee Choi, Jong-Kook Kim

Detecting fake news has received a lot of attention. Many previous methods concatenate independently encoded unimodal data, ignoring the benefits of integrated multimodal information. Also, the absence of specialized fea…

Fake News Detection