paper-with-me

홈 › Papers

Active Data Curation Effectively Distills Large-Scale Multimodal Models

2024-11-27 · CVPR 2025 1 · Vishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans, Samuel Albanie, Federico Tombari, Yongqin Xian, Alessio Tonioni, Olivier J. Hénaff

Knowledge distillation (KD) is the de facto standard for compressing large-scale models into smaller ones. Prior works have explored ever more complex KD strategies involving different objective functions, teacher-ensembles, and weight inheritance. In this work we explore an alternative, yet simple approach -- active data curation as effective distillation for contrastive multimodal pretraining. Our simple online batch selection method, ACID, outperforms strong KD baselines across various model-, data- and compute-configurations. Further, we find such an active data curation strategy to in fact be complementary to standard KD, and can be effectively combined to train highly performant inference-efficient models. Our simple and scalable pretraining framework, ACED, achieves state-of-the-art results across 27 zero-shot classification and retrieval tasks with upto 11% less inference FLOPs. We further demonstrate that our ACED models yield strong vision-encoders for training generative multimodal models in the LiT-Decoder setting, outperforming larger vision encoders for image-captioning and visual question-answering tasks.

📄 PDF Abstract BibTeX arXiv:2411.18674

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderImage CaptioningKnowledge DistillationQuestion AnsweringVisual Question Answeringzero-shot-classificationZero-Shot Learning

Similar Papers 제목 키워드 기반

Oasis: Data Curation and Assessment System for Pretraining of Large Language Models

2023-11-21 · Tong Zhou, Yubo Chen, Pengfei Cao, Kang Liu 외

Data is one of the most critical elements in building a large language model. However, existing systems either fail to customize a corpus curation pipeline or neglect to leverage comprehensive corpus assessment for itera…

Language ModelingLanguage ModellingLarge Language Model

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

2025-05-28 · Mehdi Ali, Manuel Brack, Max Lübbering, Elias Wendt 외

High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets remains limited. Existing state-of-the-ar…

Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video

2024-04-15 · Hongchi Xia, Zhi-Hao Lin, Wei-Chiu Ma, Shenlong Wang

Creating high-quality and interactive virtual environments, such as games and simulators, often involves complex and costly manual modeling processes. In this paper, we present Video2Game, a novel approach that automatic…

NeRF

Video2Game: Real-time Interactive Realistic and Browser-Compatible Environment from a Single Video

2024-01-01 · CVPR 2024 1 · Hongchi Xia, Zhi-Hao Lin, Wei-Chiu Ma, Shenlong Wang

Creating high-quality and interactive virtual environments such as games and simulators often involves complex and costly manual modeling processes. In this paper we present Video2Game a novel approach that automatic…

NeRF

Global Intelligent Content: Active Curation of Language Resources using Linked Data

2014-05-01 · LREC 2014 5 · David Lewis, Rob Brennan, Leroy Finn, Dominic Jones 외

As language resources start to become available in linked data formats, it becomes relevant to consider how linked data interoperability can play a role in active language processing workflows as well as for more static …

Information RetrievalMachine TranslationTranslation