paper-with-me

홈 › Papers

MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

2024-09-30 · Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, BoWen Zhang, Yanghao Li, Sam Dodge, Keen You, Zhen Yang, Aleksei Timofeev, Mingze Xu, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You, ZiRui Wang, Afshin Dehghan, Peter Grasch, Yinfei Yang

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systematically exploring the impact of diverse data mixtures across the entire model training lifecycle. This includes high-quality OCR data and synthetic captions for continual pre-training, as well as an optimized visual instruction-tuning data mixture for supervised fine-tuning. Our models range from 1B to 30B parameters, encompassing both dense and mixture-of-experts (MoE) variants, and demonstrate that careful data curation and training strategies can yield strong performance even at small scales (1B and 3B). Additionally, we introduce two specialized variants: MM1.5-Video, designed for video understanding, and MM1.5-UI, tailored for mobile UI understanding. Through extensive empirical studies and ablations, we provide detailed insights into the training processes and decisions that inform our final designs, offering valuable guidance for future research in MLLM development.

📄 PDF Abstract BibTeX arXiv:2409.20566

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-ExpertsOptical Character Recognition (OCR)Video UnderstandingVisual Question Answering

Similar Papers 제목 키워드 기반

MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation

2026-04-18 · Bo Li, Ningyuan Deng, Tianyu Dong, Shaobo Wang 외 arxiv

Multimodal large language models (MLLMs) have shown impressive capabilities, yet they often struggle to effectively capture the fine-grained textual information within images crucial for accurate image translation. This …

TRACER: Persistent Regularization for Robust Multimodal Finetuning

2026-05-28 · Hesam Asadollahzadeh, Feng Liu, Christopher Leckie, Sarah M. Erfani arxiv

Mainstream strategies for finetuning pretrained multimodal models often degrade out-of-distribution (OOD) robustness, a phenomenon known as catastrophic forgetting. In this paper, we develop a theoretical framework for m…

Contrastive Learning

LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation

2025-07-25 · Jingxuan Wei, Caijun Jia, Qi Chen, Yujun Cai 외 arxiv

Multimodal Machine Translation (MMT) enhances translation quality by incorporating visual context, helping to resolve textual ambiguities. While existing MMT methods perform well in bilingual settings, extending them to …

Multimodal Machine Translation

Transfer Learning with Joint Fine-Tuning for Multimodal Sentiment Analysis

2022-10-11 · Guilherme Lourenço de Toledo, Ricardo Marcondes Marcacini

Most existing methods focus on sentiment analysis of textual data. However, recently there has been a massive use of images and videos on social platforms, motivating sentiment analysis from other modalities. Current stu…

Multimodal Sentiment AnalysisSentiment AnalysisSentiment ClassificationTransfer Learning

LLaVAC: Fine-tuning LLaVA as a Multimodal Sentiment Classifier

2025-02-05 · T. Chay-intr, Y. Chen, K. Viriyayudhakorn, T. Theeramunkong

We present LLaVAC, a method for constructing a classifier for multimodal sentiment analysis. This method leverages fine-tuning of the Large Language and Vision Assistant (LLaVA) to predict sentiment labels across both im…

Multimodal Sentiment AnalysisSentiment AnalysisSentiment Classification