paper-with-me

홈 › Papers

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone

2026-05-12 · DatologyAI, :, Siddharth Joshi, Haoli Yin, Rishabh Adiga, Haakon Mongstad, Alvin Deng, Aldo Carranza, Alex Fang, Amro Abbas, Anshuman Suri, Brett Larsen, Daniel Zayas, Darren Teh, David Schwab, Diego Kiner, Fan Pan, Jack Urbanek, Jason Lee, Jason Telanoff, Josh Wills, Kaleigh Mentzer, Luke Merrick, Maximilian Böther, Parth Doshi, Paul Burstein, Pratyush Maini, Ties Robroek, Tony Jiang, Vidhi Jain, Vineeth Dorna, Zhengping Wang, Bogdan Gaza, Ari Morcos, Matthew Leavitt arxiv

Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less established. We ask how far data curation alone can take VLM performance, holding architecture, training recipe, and compute fixed and varying only the training data. Our pipeline, applied to the MAmmoTH-VL single-image subset, lifts performance by +11.7pp on average across 20 public VLM benchmarks (spanning grounding, VQA, OCR/documents, captioning, spatial/3D, counting, charts, math, brand-ID, and multi-image reasoning) and by +11.3pp on average across all nine capability axes of DatBench, our high-fidelity VLM eval suite. At 2B, our curated model surpasses InternVL3.5-2B by 9.9pp at ~17x less training compute and closes the gap to Qwen3-VL-2B to within 1.8pp at ~87x less compute, from pretraining alone. Beyond accuracy, curation delivers four further properties: (1) Reliability: per-capability std across training seeds drops by ~67% and the lift survives a 4k-to-16k context-length sweep; (2) OOD generalization: the 9-eval OOD average rises by +7.2pp, and multi-image BLINK rises by +3.09pp despite single-image-only training, with Visual Correspondence gaining +11.8pp; (3) Behavioral gains beyond benchmarks: across ~1,100 open-ended queries the curated 2B is more honest and more specific than the matched-compute baseline, and more concise and less refusal-prone than a frontier 2B reference; (4) Pareto-dominance on inference cost: at every scale (1B, 2B, 4B) the curated model raises accuracy while lowering response FLOPs vs. the matched-compute baseline, and the curated 4B matches near-frontier accuracy at 3.3x lower response FLOPs than Qwen3-VL-4B. Data curation is a high-leverage tool for building better VLMs, reaching near-frontier accuracy at up to ~150x less training compute.

📄 PDF Abstract BibTeX arXiv:2605.11405

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models

2026-02-19 · Dhruba Ghosh, Yuhui Zhang, Ludwig Schmidt arxiv

Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and multimodal dialogue. These improvements are…

Visual Question AnsweringImage ClassificationVisual Reasoning

An Introduction to Vision-Language Modeling

2024-05-27 · Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay, Alexander C. Li 외

Following the recent popularity of Large Language Models (LLMs), several attempts have been made to extend them to the visual domain. From having a visual assistant that could guide us through unfamiliar environments to …

Language ModelingLanguage Modelling

ChromaCorrect: Prescription Correction in Virtual Reality Headsets through Perceptual Guidance

2022-12-08 · Ahmet Güzel, Jeanne Beyazian, PRANEETH CHAKRAVARTHULA, Kaan Akşit

A large portion of today's world population suffer from vision impairments and wear prescription eyeglasses. However, eyeglasses causes additional bulk and discomfort when used with augmented and virtual reality headsets…

Attention Guided Alignment in Efficient Vision-Language Models

2025-11-21 · Shweta Mahajan, Hoang Le, Hyojin Park, Farzad Farhadzadeh 외 arxiv

Large Vision-Language Models (VLMs) rely on effective multimodal alignment between pre-trained vision encoders and Large Language Models (LLMs) to integrate visual and textual information. This paper presents a comprehen…

Visual Grounding

Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning

2022-07-15 · Shibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang 외

The development of CLIP [Radford et al., 2021] has sparked a debate on whether language supervision can result in vision models with more transferable representations than traditional image-only methods. Our work studies…

DescriptiveRepresentation Learning