paper-with-me

홈 › Papers

LinMU: Multimodal Understanding Made Linear

2026-01-04 · Hongjie Wang, Niraj K. Jha arxiv

Modern Vision-Language Models (VLMs) achieve impressive performance but are limited by the quadratic complexity of self-attention, which prevents their deployment on edge devices and makes their understanding of high-resolution images and long-context videos prohibitively expensive. To address this challenge, we introduce LinMU (Linear-complexity Multimodal Understanding), a VLM design that achieves linear complexity for the language model decoder without using any quadratic-complexity modules while maintaining the performance of global-attention-based VLMs. LinMU replaces every self-attention layer in the language model decoder with an M-MATE block: a dual-branch module that combines a bidirectional state-space model for global context (Flex-MA branch) with localized Swin-style window attention (Local-Swin branch) for adjacent correlations. To transform a pre-trained VLM into the LinMU architecture, we propose a three-stage distillation framework that (i) initializes both branches with self-attention weights and trains the Flex-MA branch alone, (ii) unfreezes the Local-Swin branch and fine-tunes it jointly with the Flex-MA branch, and (iii) unfreezes the remaining blocks and fine-tunes them using LoRA adapters, while regressing on hidden states and token-level logits of the frozen VLM teacher. On MMMU, TextVQA, LongVideoBench, Video-MME, and other benchmarks, LinMU matches the performance of teacher models, yet reduces Time-To-First-Token (TTFT) by up to 2.7$\times$ and improves token throughput by up to 9.0$\times$ on minute-length videos. Ablations confirm the importance of each distillation stage and the necessity of the two branches of the M-MATE block. The proposed framework demonstrates that state-of-the-art multimodal reasoning can be achieved without quadratic attention, thus opening up avenues for long-context VLMs that can deal with high-resolution images and long videos.

📄 PDF Abstract BibTeX arXiv:2601.01322

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation

2025-02-17 · Ling Yang, Xinchen Zhang, Ye Tian, Chenming Shang 외

The remarkable success of the autoregressive paradigm has made significant advancement in Multimodal Large Language Models (MLLMs), with powerful models like Show-o, Transfusion and Emu3 achieving notable progress in uni…

ChemMLLM: Chemical Multimodal Large Language Model

2025-05-22 · Qian Tan, Dongzhan Zhou, Peng Xia, Wanhao Liu 외

Multimodal large language models (MLLMs) have made impressive progress in many applications in recent years. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill …

Language ModelingLanguage ModellingLarge Language Modelmodel+1

AIN: The Arabic INclusive Large Multimodal Model

2025-01-31 · Ahmed Heakl, Sara Ghaboura, Omkar Thawkar, Fahad Shahbaz Khan 외

Amid the swift progress of large language models (LLMs) and their evolution into large multimodal models (LMMs), significant strides have been made in high-resource languages such as English and Chinese. While Arabic LLM…

document understandingmodelVideo Understanding

UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

2023-08-19 · Hao Feng, Zijian Wang, Jingqun Tang, Jinghui Lu 외

In the era of Large Language Models (LLMs), tremendous strides have been made in the field of multimodal understanding. However, existing advanced algorithms are limited to effectively utilizing the immense representatio…

Instruction FollowingText DetectionWorld Knowledge

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

2026-03-06 · Lijiang Li, Zuwei Long, Yunhang Shen, Heting Gao 외 arxiv

While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective an…

Image Generation