ICU: Conquering Language Barriers in Vision-and-Language Modeling by Dividing the Tasks into Image Captioning and Language Understanding
Most multilingual vision-and-language (V&L) research aims to accomplish multilingual and multimodal capabilities within one model. However, the scarcity of multilingual captions for images has hindered the development. To overcome this obstacle, we propose ICU, Image Caption Understanding, which divides a V&L task into two stages: a V&L model performs image captioning in English, and a multilingual language model (mLM), in turn, takes the caption as the alt text and performs cross-lingual language understanding. The burden of multilingual processing is lifted off V&L model and placed on mLM. Since the multilingual text data is relatively of higher abundance and quality, ICU can facilitate the conquering of language barriers for V&L models. In experiments on two tasks across 9 languages in the IGLUE benchmark, we show that ICU can achieve new state-of-the-art results for five languages, and comparable results for the rest.
Code (1)
Tasks
Image CaptioningLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
Large Multimodal Models (LMMs) have achieved significant breakthroughs in various vision-language and vision-centric tasks based on auto-regressive modeling. However, these models typically focus on either vision-centric…
Instruction Followingobject-detectionObject DetectionQuestion Answering+5The Competitiveness Analysis of the European Language Technology Market
This paper presents the key results of a study on the global competitiveness of the European Language Technology market for three areas {--} Machine Translation, speech technology, and cross-lingual search. EU competitiv…
Machine TranslationTranslationBreaking the Barriers: Video Vision Transformers for Word-Level Sign Language Recognition
Sign language is a fundamental means of communication for the deaf and hard-of-hearing (DHH) community, enabling nuanced expression through gestures, facial expressions, and body movements. Despite its critical role in f…
Gesture RecognitionSign Language RecognitionOpen-Vocabulary Point-Cloud Object Detection without 3D Annotation
The goal of open-vocabulary detection is to identify novel objects based on arbitrary textual descriptions. In this paper, we address open-vocabulary 3D point-cloud detection by a dividing-and-conquering strategy, which …
3D Object Detection3D Open-Vocabulary Object DetectionCloud DetectionContrastive Learning+3ENSEMBITS: an alphabet of protein conformational ensembles
Protein structure tokenizers (PSTs) are workhorses in protein language modeling, function prediction, and evolutionary analysis. However, existing PSTs only capture local geometry of static structures, and miss the corre…