paper-with-me

Papers

MammothModa: Multi-Modal Large Language Model

2024-06-26 · Qi She, Junwen Pan, Xin Wan, Rui Zhang, Dawei Lu, Kai Huang

In this report, we introduce MammothModa, yet another multi-modal large language model (MLLM) designed to achieve state-of-the-art performance starting from an elementary baseline. We focus on three key design insights: (i) Integrating Visual Capabilities while Maintaining Complex Language Understanding: In addition to the vision encoder, we incorporated the Visual Attention Experts into the LLM to enhance its visual capabilities. (ii) Extending Context Window for High-Resolution and Long-Duration Visual Feature: We explore the Visual Merger Module to effectively reduce the token number of high-resolution images and incorporated frame position ids to avoid position interpolation. (iii) High-Quality Bilingual Datasets: We meticulously curated and filtered a high-quality bilingual multimodal dataset to reduce visual hallucinations. With above recipe we build MammothModa that consistently outperforms the state-of-the-art models, e.g., LLaVA-series, across main real-world visual language benchmarks without bells and whistles.

📄 PDF Abstract BibTeX arXiv:2406.18193

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelmodelPosition

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation

2025-11-23 · Tao Shen, Xin Wan, Taicai Chen, Rui Zhang 외 arxiv

Unified multimodal models aim to integrate understanding and generation within a single framework, yet bridging the gap between discrete semantic reasoning and high-fidelity visual synthesis remains challenging. We prese…

Reinforcement Learning

CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French

2020-11-01 · EMNLP 2020 11 · AmirAli Bagher Zadeh, Yansheng Cao, Simon Hessner, Paul Pu Liang 외

Modeling multimodal language is a core research area in natural language processing. While languages such as English have relatively large multimodal language resources, other widely spoken languages across the globe hav…

Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

2023-08-23 · Jinyi Hu, Yuan YAO, Chongyi Wang, Shan Wang 외

Recently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English, leaving other languages largely behind…

Image GenerationImage to textLanguage ModelingLanguage Modelling+3

OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects

2024-10-02 · Wenmo Qiu, Xinhan Di

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multimodal models fail to provide satisfactory results in describing occluded o…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

2023-05-07 · Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang 외

Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual language models. We attribute this to the…

AttributeInstruction FollowingLanguage ModellingLarge Language Model+2