paper-with-me

홈 › Papers

SAIL-VL2 Technical Report

2025-09-17 · Weijie Yin, Yongjie Ye, Fangxun Shu, Yue Liao, Zijian Kang, Hongyuan Dong, Haiyang Yu, Dingkang Yang, Jiacong Wang, Han Wang, Wenzhuo Liu, Xiao Liang, Shuicheng Yan, Chao Feng arxiv

We introduce SAIL-VL2, an open-suite vision-language foundation model (LVM) for comprehensive multimodal understanding and reasoning. As the successor to SAIL-VL, SAIL-VL2 achieves state-of-the-art performance at the 2B and 8B parameter scales across diverse image and video benchmarks, demonstrating strong capabilities from fine-grained perception to complex reasoning. Its effectiveness is driven by three core innovations. First, a large-scale data curation pipeline with scoring and filtering strategies enhances both quality and distribution across captioning, OCR, QA, and video data, improving training efficiency. Second, a progressive training framework begins with a powerful pre-trained vision encoder (SAIL-ViT), advances through multimodal pre-training, and culminates in a thinking-fusion SFT-RL hybrid paradigm that systematically strengthens model capabilities. Third, architectural advances extend beyond dense LLMs to efficient sparse Mixture-of-Experts (MoE) designs. With these contributions, SAIL-VL2 demonstrates competitive performance across 106 datasets and achieves state-of-the-art results on challenging reasoning benchmarks such as MMMU and MathVista. Furthermore, on the OpenCompass leaderboard, SAIL-VL2-2B ranks first among officially released open-source models under the 4B parameter scale, while serving as an efficient and extensible foundation for the open-source multimodal community.

📄 PDF Abstract BibTeX arXiv:2509.14033

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Measurement Models For Sailboats Price vs. Features And Regional Areas

2023-09-26 · Jiaqi Weng, Chunlin Feng, Yihan Shao

In this study, we investigated the relationship between sailboat technical specifications and their prices, as well as regional pricing influences. Utilizing a dataset encompassing characteristics like length, beam, draf…

FL-Sailer: Efficient and Privacy-Preserving Federated Learning for Scalable Single-Cell Epigenetic Data Analysis via Adaptive Sampling

2026-05-06 · Guangyi Zhang, Yi Dai, Yiyun He, Junhao Liu arxiv

Single-cell ATAC-seq (scATAC-seq) enables high-resolution mapping of chromatin accessibility, yet privacy regulations and data size constraints hinder multi-institutional sharing. Federated learning (FL) offers a privacy…

Federated Learning

SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model

2025-10-14 · Lin Lin, Jiefeng Long, Zhihe Wan, Yuchi Wang 외 arxiv

Multimodal embedding models aim to yield informative unified representations that empower diverse cross-modal tasks. Despite promising developments in the evolution from CLIP-based dual-tower architectures to large visio…

Representation Learning

Benchmarking Large Multimodal Models against Common Corruptions

2024-01-22 · Jiawei Zhang, Tianyu Pang, Chao Du, Yi Ren 외

This technical report aims to fill a deficiency in the assessment of large multimodal models (LMMs) by specifically examining the self-consistency of their outputs when subjected to common corruptions. We investigate the…

BenchmarkingImage to textSpeech-to-Texttext-to-speech+1

Sailor: Open Language Models for South-East Asia

2024-04-04 · Longxu Dou, Qian Liu, Guangtao Zeng, Jia Guo 외

We present Sailor, a family of open language models ranging from 0.5B to 7B parameters, tailored for South-East Asian (SEA) languages. These models are continually pre-trained from Qwen1.5, a great language model for mul…

Language ModelingLanguage ModellingQuestion AnsweringReading Comprehension