paper-with-me

홈 › Papers

ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

2024-12-09 · Jieyu Zhang, Le Xue, Linxin Song, Jun Wang, Weikai Huang, Manli Shu, An Yan, Zixian Ma, Juan Carlos Niebles, Silvio Savarese, Caiming Xiong, Zeyuan Chen, Ranjay Krishna, ran Xu

With the rise of multimodal applications, instruction data has become critical for training multimodal language models capable of understanding complex image-based queries. Existing practices rely on powerful but costly large language models (LLMs) or multimodal language models (MLMs) to produce instruction data. These are often prone to hallucinations, licensing issues and the generation process is often hard to scale and interpret. In this work, we present a programmatic approach that employs scene graphs as symbolic representations of images and human-written programs to systematically synthesize vision-centric instruction data. Our approach ensures the interpretability and controllability of the data generation process and scales efficiently while maintaining factual accuracy. By implementing a suite of 24 single-image, 14 multi-image instruction generators, and a scene graph generation pipeline, we build a scalable, cost-effective system: ProVision which produces diverse question-answer pairs concerning objects, attributes, relations, depth, etc., for any given image. Applied to Visual Genome and DataComp datasets, we generate over 10 million instruction data points, ProVision-10M, and leverage them in both pretraining and instruction tuning stages of MLMs. When adopted in the instruction tuning stage, our single-image instruction data yields up to a 7% improvement on the 2D split and 8% on the 3D split of CVBench, along with a 3% increase in performance on QBench2, RealWorldQA, and MMMU. Our multi-image instruction data leads to an 8% improvement on Mantis-Eval. Incorporation of our data in both pre-training and fine-tuning stages of xGen-MM-4B leads to an averaged improvement of 1.6% across 11 benchmarks.

📄 PDF Abstract BibTeX arXiv:2412.07012

Code (1)

jieyuz2/provision 공식 구현 pytorch

Tasks

Graph GenerationScene Graph GenerationVisual Question Answering

Similar Papers 제목 키워드 기반

User-Centric Communication Service Provision for Edge-Assisted Mobile Augmented Reality

2025-09-30 · Conghao Zhou, Jie Gao, Shisheng Hu, Nan Cheng 외 arxiv

Future 6G networks are envisioned to facilitate edge-assisted mobile augmented reality (MAR) via strengthening the collaboration between MAR devices and edge servers. In order to provide immersive user experiences, MAR d…

Pose Tracking

AI-based Resource Allocation: Reinforcement Learning for Adaptive Auto-scaling in Serverless Environments

2020-05-29 · Lucia Schuler, Somaya Jamil, Niklas Kühl

Serverless computing has emerged as a compelling new paradigm of cloud computing models in recent years. It promises the user services at large scale and low cost while eliminating the need for infrastructure management.…

Cloud ComputingManagementreinforcement-learningReinforcement Learning (RL)

Statistical QoS Provision in Business-Centric Networks

2024-08-28 · Chang Wu, Yuang Chen, Hancheng Lu

More refined resource management and Quality of Service (QoS) provisioning is a critical goal of wireless communication technologies. In this paper, we propose a novel Business-Centric Network (BCN) aimed at enabling sca…

Deep Reinforcement Learning

Wasserstein Adversarial Transformer for Cloud Workload Prediction

2022-03-12 · Shivani Arbat, Vinodh Kumaran Jayakumar, Jaewoo Lee, Wei Wang 외

Predictive Virtual Machine (VM) auto-scaling is a promising technique to optimize cloud applications operating costs and performance. Understanding the job arrival rate is crucial for accurately predicting future changes…

PredictionTime SeriesTime Series AnalysisTime Series Forecasting

Accelerating Serverless Computing by Harvesting Idle Resources

2021-08-28 · Hanfei Yu, Hao Wang, Jian Li, Xu Yuan 외

Serverless computing automates fine-grained resource scaling and simplifies the development and deployment of online services with stateless functions. However, it is still non-trivial for users to allocate appropriate r…

Deep Reinforcement Learning