paper-with-me

홈 › Papers

Scaling Laws for Native Multimodal Models Scaling Laws for Native Multimodal Models

2025-04-10 · Mustafa Shukor, Enrico Fini, Victor Guilherme Turrisi da Costa, Matthieu Cord, Joshua Susskind, Alaaeldin El-Nouby

Building general-purpose models that can effectively perceive the world through multimodal signals has been a long-standing goal. Current approaches involve integrating separately pre-trained components, such as connecting vision encoders to LLMs and continuing multimodal training. While such approaches exhibit remarkable sample efficiency, it remains an open question whether such late-fusion architectures are inherently superior. In this work, we revisit the architectural design of native multimodal models (NMMs)--those trained from the ground up on all modalities--and conduct an extensive scaling laws study, spanning 457 trained models with different architectures and training mixtures. Our investigation reveals no inherent advantage to late-fusion architectures over early-fusion ones, which do not rely on image encoders. On the contrary, early-fusion exhibits stronger performance at lower parameter counts, is more efficient to train, and is easier to deploy. Motivated by the strong performance of the early-fusion architectures, we show that incorporating Mixture of Experts (MoEs) allows for models that learn modality-specific weights, significantly enhancing performance.

📄 PDF Abstract BibTeX arXiv:2504.07951

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-Experts

Similar Papers 제목 키워드 기반

Scaling Laws for Optimal Data Mixtures

2025-07-12 · Mustafa Shukor, Louis Bethune, Dan Busbridge, David Grangier 외 arxiv

Large foundation models are typically trained on data from multiple domains, with the data mixture--the proportion of each domain used--playing a critical role in model performance. The standard approach to selecting thi…

Scaling Laws for Autoregressive Generative Modeling

2020-10-28 · Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen 외

We identify empirical scaling laws for the cross-entropy loss in four domains: generative image modeling, video modeling, multimodal image$\leftrightarrow$text models, and mathematical problem solving. In all cases autor…

Mathematical Problem-Solving

Scaling laws in wearable human activity recognition

2025-02-05 · Tom Hoddes, Alex Bijamov, Saket Joshi, Daniel Roggen 외

Many deep architectures and self-supervised pre-training techniques have been proposed for human activity recognition (HAR) from wearable multimodal sensors. Scaling laws have the potential to help move towards more prin…

Activity RecognitionHuman Activity Recognition

On the origin of neural scaling laws: from random graphs to natural language

2026-01-15 · Maissam Barkeshli, Alberto Alfarano, Andrey Gromov arxiv

Scaling laws have played a major role in the modern AI revolution, providing practitioners predictive power over how the model performance will improve with increasing data, compute, and number of model parameters. This …

Scaling Laws for Discriminative Speech Recognition Rescoring Models

2023-06-27 · Yile Gu, Prashanth Gurunath Shivakumar, Jari Kolehmainen, Ankur Gandhe 외

Recent studies have found that model performance has a smooth power-law relationship, or scaling laws, with training data and model size, for a wide range of problems. These scaling laws allow one to choose nearly optima…

speech-recognitionSpeech Recognition