paper-with-me

Papers

Scaling Law Hypothesis for Multimodal Model

2024-09-10 · Qingyun Sun, Zhen Guo, PIN AI Team

We propose a scaling law hypothesis for multimodal models processing text, audio, images, and video within a shared token and embedding space. Our framework predicts model performance based on modality-specific compression and tokenization efficiency, extending established scaling laws from text-based decoder models to mixed-modality systems. We explore whether leveraging more training data in multiple modalities can reduce the size of the multimodal model, enabling efficient deployment on resource-constrained devices.

📄 PDF Abstract BibTeX arXiv:2409.06754

Code (0)

등록된 구현이 없습니다.

Tasks

Decodermodel

Similar Papers 제목 키워드 기반

From Zipf's Law to Neural Scaling through Heaps' Law and Hilberg's Hypothesis

2025-12-15 · Łukasz Dębowski arxiv

We inspect the deductive connection between the neural scaling law and Zipf's law -- two statements discussed in machine learning and quantitative linguistics. The neural scaling law describes how the cross entropy rate …

Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition

2025-03-16 · Xiaoying Zhang, Da Peng, YiPeng Zhang, Zonghao Guo 외

Recent progress in (multimodal) large language models ((M)LLMs) has shifted focus from pre-training to inference-time compute scaling and post-training optimization, driven by concerns over limited high-quality real-worl…

Caption GenerationImage CaptioningMultimodal ReasoningSelf-Learning

Multiple power laws and scaling relation in exploratory locomotion of the snail Tegula nigerrima

2024-10-28 · Katsushi Kagaya, Tomoyuki Nakano, Ryo Nakayama

One of goals in soft robotics is to achive spontaneous behavior like real organisms. To gain a clue to achieve this, we examined the long (16-hour) spontaneous exploratory locomotion of snails. The active forager snail, …

Relation

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

2026-04-22 · Karan Goyal arxiv

The rapid proliferation of Vision-Language Models (VLMs) is often framed as enabling unified multimodal knowledge discovery but rests on an under-examined assumption: that current VLMs faithfully synthesise multimodal da…

Multimodal Reasoning

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct

2026-06-22 · Haoling Li, Kai Zheng, Jie Wu, Can Xu 외 arxiv

Scaling reinforcement learning for visual mathematical reasoning requires more than generating harder questions: as data volume grows, the reward labels themselves must remain reliable. Yet existing data pipelines scale …

Reinforcement LearningMathematical Reasoning