paper-with-me

홈 › Papers

Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models

2025-09-27 · Junjie Li, Ziao Wang, Jianghong Ma, Xiaofeng Zhang arxiv

Large vision-language models (VLMs) achieve strong benchmark performance, but controlling their behavior through instruction tuning remains difficult. Reducing the budget of instruction tuning dataset often causes regressions, as heuristic strategies treat models as black boxes and overlook the latent capabilities that govern learning. We introduce Capability-Attributed Data Curation (CADC), a framework that shifts curation from task-specific heuristics to intrinsic capability analysis. CADC discovers intrinsic capabilities in an unsupervised manner from gradient-based learning trajectories, attributes training data to these capabilities via influence estimation, and curates capability-aware curricula through balanced selection and staged sequencing. This transforms black-box instruction tuning into a controllable, capability-driven process. With as little as 5% of the original data, CADC surpasses full-data training on multimodal benchmarks. These results validate intrinsic capabilities as the fundamental building blocks of model learning and establish CADC as a principle paradigm for instruction data curation.

📄 PDF Abstract BibTeX arXiv:2510.00040

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution

2026-05-25 · Zehao Wang, Yihan Zeng, Zidong Gong, Yuanfan Guo 외 arxiv

Post-training via Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) is crucial for enhancing reasoning in Multimodal Large Language Models (MLLMs), yet existing paradigms often reach a performance bottleneck d…

Reinforcement LearningMultimodal Reasoning

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs

2026-04-12 · Yu Li, Xiaoran Shang, Qizhi Pei, Yun Zhu 외 arxiv

Post-training data plays a pivotal role in shaping the capabilities of Large Language Models (LLMs), yet datasets are often treated as isolated artifacts, overlooking the systemic connections that underlie their evolutio…

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

2025-06-09 · Vishaal Udandarao, Mehdi Cherti, Shyamgopal Karthik, Jenia Jitsev 외

We investigate 17 benchmarks (e.g. SugarCREPE, VALSE) commonly used for measuring compositional understanding capabilities of vision-language models (VLMs). We scrutinize design choices in their construction, including d…

Language ModelingLanguage Modelling

Best practices for the manual curation of Intrinsically Disordered Proteins in DisProt

2023-10-25 · Federica Quaglia, Anastasia Chasapi, Maria Victoria Nugnes, Maria Cristina Aspromonte 외

The DisProt database is a significant resource containing manually curated data on experimentally validated intrinsically disordered proteins (IDPs) and regions (IDRs) from the literature. Developed in 2005, its primary …

AI4D - African Language Dataset Challenge

2020-07-01 · WS 2020 7 · Kathleen Siminyu, Sackey Freshia

As language and speech technologies become more advanced, the lack of fundamental digital resources for African languages, such as data, spell checkers and PoS taggers, means that the digital divide between these languag…

POS