paper-with-me

홈 › Papers

NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference

2026-08-31 · Aurélien Lac, Tony Wu hf

Multimodal models often build on architectures designed for generative vision-language modeling, typically combining separately pretrained vision encoders with causal language models. Visual document retrievers such as ColPali repurpose these models as encoders, carrying over the parameter and compute overhead of a VLM for a non-generative task. We introduce NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder. Both models are pretrained from scratch with a masked discrete-diffusion text objective, conditioned on visible image patches for multimodal examples. Both support a 16,384-token context, enough to encode up to two standard 4K UHD images. To demonstrate its downstream capabilities, we fine-tune NeoMME with jointly trained dense and late-interaction heads. On the ViDoRe v3 benchmark, the resulting NeoMME-Retriever 260M outperforms all evaluated models strictly below 800M parameters with 0.523 nDCG@10, while NeoMME-Retriever 800M reaches 0.556. At a matched 2048x2048 image input size on an NVIDIA L40S, NeoMME-260M encodes pages with about 2x the throughput of ColModernVBERT. Hierarchical token pooling and asymmetric quantization compress late-interaction multimodal document embeddings by 255x while preserving over 95% of baseline nDCG@10. We contribute NeoMME to Hugging Face Transformers and release the pretrained backbone and retrieval-compatible checkpoints under Apache 2.0 at https://hf.co/collections/Hcompany/neomme.

📄 PDF Abstract BibTeX arXiv:2609.01657

Code (8)

Valiant-Cat/hfpaper
arxivsub/arXivSub_daily_arxiv ★ 4
tonywu71/neomme-retriever-demo
🤗 Hcompany/NeoMME-260M ★ 5
🤗 Hcompany/NeoMME-260M-Retriever ★ 4
🤗 Hcompany/NeoMME-260M-Retriever-ST-dense ★ 4
🤗 Hcompany/NeoMME-260M-Retriever-ST-late ★ 5
🤗 Hcompany/NeoMME-800M-Retriever ★ 4

Similar Papers 제목 키워드 기반

Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

2024-02-27 · Duarte M. Alves, José Pombal, Nuno M. Guerreiro, Pedro H. Martins 외

While general-purpose large language models (LLMs) demonstrate proficiency on multiple tasks within the domain of translation, approaches based on open LLMs are competitive only when specializing on a single task. In thi…

Language ModelingLanguage ModellingLarge Language ModelTranslation

TowerVision: Understanding and Improving Multilinguality in Vision-Language Models

2025-10-22 · André G. Viveiros, Patrick Fernandes, Saul Santos, Sonal Sannigrahi 외 arxiv

Despite significant advances in vision-language models (VLMs), most existing work follows an English-centric design process, limiting their effectiveness in multilingual settings. In this work, we provide a comprehensive…

The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model

2024-12-10 · Jiawei Chen, Wentao Chen, Jing Su, Jingjing Xu 외

Large language models (LLMs) have shown significant multilingual capabilities. However, the mechanisms underlying the development of these capabilities during pre-training are not well understood. In this paper, we use c…

Language ModelingLanguage ModellingLarge Language Model

CLIPPO: Image-and-Language Understanding from Pixels Only

2022-12-15 · CVPR 2023 1 · Michael Tschannen, Basil Mustafa, Neil Houlsby

Multimodal models are becoming increasingly effective, in part due to unified components, such as the Transformer architecture. However, multimodal models still often consist of many task- and modality-specific pieces an…

Contrastive Learningimage-classificationImage ClassificationLanguage Modelling+7

Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context

2026-06-25 · Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary 외 arxiv

Diffusion language models offer a promising alternative to autoregressive models due to their potential for parallel and iterative generation. However, existing approaches use a single network for both context representa…