paper-with-me

홈 › Papers

Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

2023-04-14 · NeurIPS 2023 11 · Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, Yejin Choi

In-context vision and language models like Flamingo support arbitrarily interleaved sequences of images and text as input. This format not only enables few-shot learning via interleaving independent supervised (image, text) examples, but also, more complex prompts involving interaction between images, e.g., "What do image A and image B have in common?" To support this interface, pretraining occurs over web corpora that similarly contain interleaved images+text. To date, however, large-scale data of this form have not been publicly available. We release Multimodal C4, an augmentation of the popular text-only C4 corpus with images interleaved. We use a linear assignment algorithm to place images into longer bodies of text using CLIP features, a process that we show outperforms alternatives. Multimodal C4 spans everyday topics like cooking, travel, technology, etc. A manual inspection of a random sample of documents shows that a vast majority (88%) of images are topically relevant, and that linear assignment frequently selects individual sentences specifically well-aligned with each image (80%). After filtering NSFW images, ads, etc., the resulting corpus consists of 101.2M documents with 571M images interleaved in 43B English tokens.

📄 PDF Abstract BibTeX arXiv:2304.06939

Code (1)

allenai/mmc4 공식 구현 jax

Tasks

Few-Shot Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

2024-06-12 · Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang 외

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studie…

In-Context Learning

OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

2023-06-21 · NeurIPS 2023 11 · Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman 외

Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models hav…

MMR total

MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens

2024-06-17 · Anas Awadalla, Le Xue, Oscar Lo, Manli Shu 외

Multimodal interleaved datasets featuring free-form interleaved sequences of images and text are crucial for training frontier large multimodal models (LMMs). Despite the rapid progression of open-source LMMs, there rema…

Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl

2017-10-04 · LREC 2018 5 · Alexander Panchenko, Eugen Ruppert, Stefano Faralli, Simone Paolo Ponzetto 외

We present DepCC, the largest-to-date linguistically analyzed corpus in English including 365 million documents, composed of 252 billion tokens and 7.5 billion of named entity occurrences in 14.3 billion sentences from a…

Open Information ExtractionQuestion AnsweringWord Embeddings

Construction and Applications of Billion-Scale Pre-Trained Multimodal Business Knowledge Graph

2022-09-30 · Shumin Deng, Chengming Wang, Zhoubo Li, Ningyu Zhang 외

Business Knowledge Graphs (KGs) are important to many enterprises today, providing factual knowledge and structured data that steer many products and make them more intelligent. Despite their promising benefits, building…

Knowledge Graphs