paper-with-me

홈 › Papers

ICU: Conquering Language Barriers in Vision-and-Language Modeling by Dividing the Tasks into Image Captioning and Language Understanding

2023-10-19 · Guojun Wu

Most multilingual vision-and-language (V&L) research aims to accomplish multilingual and multimodal capabilities within one model. However, the scarcity of multilingual captions for images has hindered the development. To overcome this obstacle, we propose ICU, Image Caption Understanding, which divides a V&L task into two stages: a V&L model performs image captioning in English, and a multilingual language model (mLM), in turn, takes the caption as the alt text and performs cross-lingual language understanding. The burden of multilingual processing is lifted off V&L model and placed on mLM. Since the multilingual text data is relatively of higher abundance and quality, ICU can facilitate the conquering of language barriers for V&L models. In experiments on two tasks across 9 languages in the IGLUE benchmark, we show that ICU can achieve new state-of-the-art results for five languages, and comparable results for the rest.

📄 PDF Abstract BibTeX arXiv:2310.12531

Code (1)

gjwubyron/icu 공식 구현 pytorch

Tasks

Image CaptioningLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models

2024-10-21 · Yufei Zhan, Hongyin Zhao, Yousong Zhu, Fan Yang 외

Large Multimodal Models (LMMs) have achieved significant breakthroughs in various vision-language and vision-centric tasks based on auto-regressive modeling. However, these models typically focus on either vision-centric…

Instruction Followingobject-detectionObject DetectionQuestion Answering+5

The Competitiveness Analysis of the European Language Technology Market

2020-05-01 · LREC 2020 5 · Andrejs Vasi{\c{l}}jevs, Inguna Skadi{\c{n}}a, Indra Samite, Kaspars Kauli{\c{n}}{\v{s}} 외

This paper presents the key results of a study on the global competitiveness of the European Language Technology market for three areas {--} Machine Translation, speech technology, and cross-lingual search. EU competitiv…

Machine TranslationTranslation

Breaking the Barriers: Video Vision Transformers for Word-Level Sign Language Recognition

2025-04-10 · Alexander Brettmann, Jakob Grävinghoff, Marlene Rüschoff, Marie Westhues

Sign language is a fundamental means of communication for the deaf and hard-of-hearing (DHH) community, enabling nuanced expression through gestures, facial expressions, and body movements. Despite its critical role in f…

Gesture RecognitionSign Language Recognition

Open-Vocabulary Point-Cloud Object Detection without 3D Annotation

2023-04-03 · CVPR 2023 1 · Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Xiaodong Xie 외

The goal of open-vocabulary detection is to identify novel objects based on arbitrary textual descriptions. In this paper, we address open-vocabulary 3D point-cloud detection by a dividing-and-conquering strategy, which …

3D Object Detection3D Open-Vocabulary Object DetectionCloud DetectionContrastive Learning+3

ENSEMBITS: an alphabet of protein conformational ensembles

2026-05-13 · Kaiwen Shi, Carlos Oliver arxiv

Protein structure tokenizers (PSTs) are workhorses in protein language modeling, function prediction, and evolutionary analysis. However, existing PSTs only capture local geometry of static structures, and miss the corre…