WIT
Wikipedia-based Image Text
홈페이지 · 논문 79편
Wikipedia-based Image Text (WIT) Dataset is a large multimodal multilingual dataset. WIT is composed of a curated set of 37.6 million entity rich image-text examples with 11.5 million unique images across 108 Wikipedia languages. Its size enables WIT to be used as a pretraining dataset for multimodal machine learning models. Key Advantages A few unique advantages of WIT: - The largest multimodal dataset (time of this writing) by the number of image-text examples. - A massively multilingual (first of its kind) with coverage for over 100+ languages. - A collection of diverse set of concepts and real world entities. - Brings forth challenging real-world test sets.
ImagesTexts Multilingual벤치마크
Image Retrieval on WIT
결과 2개