paper-with-me

홈 › Papers

Globetrotter: Connecting Languages by Connecting Images

2020-12-08 · CVPR 2022 1 · Dídac Surís, Dave Epstein, Carl Vondrick

Machine translation between many languages at once is highly challenging, since training with ground truth requires supervision between all language pairs, which is difficult to obtain. Our key insight is that, while languages may vary drastically, the underlying visual appearance of the world remains consistent. We introduce a method that uses visual observations to bridge the gap between languages, rather than relying on parallel corpora or topological properties of the representations. We train a model that aligns segments of text from different languages if and only if the images associated with them are similar and each image in turn is well-aligned with its textual description. We train our model from scratch on a new dataset of text in over fifty languages with accompanying images. Experiments show that our method outperforms previous work on unsupervised word and sentence translation using retrieval. Code, models and data are available on globetrotter.cs.columbia.edu.

📄 PDF Abstract BibTeX arXiv:2012.04631

Code (1)

cvlab-columbia/globetrotter pytorch

Tasks

Machine TranslationRetrievalSentenceTranslation

Similar Papers 제목 키워드 기반

Connecting Ideas in 'Lower-Resource' Scenarios: NLP for National Varieties, Creoles and Other Low-resource Scenarios

2024-09-19 · Aditya Joshi, Diptesh Kanojia, Heather Lent, Hour Kaing 외

Despite excellent results on benchmarks over a small subset of languages, large language models struggle to process text from languages situated in `lower-resource' scenarios such as dialects/sociolects (national or soci…

The Multilingual Anonymisation Toolkit for Public Administrations (MAPA) Project

2020-11-01 · EAMT 2020 11 · Ēriks Ajausks, Victoria Arranz, Laurent Bié, Aleix Cerdà-i-Cucó 외

We describe the MAPA project, funded under the Connecting Europe Facility programme, whose goal is the development of an open-source de-identification toolkit for all official European Union languages. It will be develop…

De-identification

Converting a Common Document Scanner to a Multispectral Scanner

2019-04-17 · Zohaib Khan, Faisal Shafait, Ajmal Mian

We propose the construction of a prototype scanner designed to capture multispectral images of documents. A standard sheet-feed scanner is modified by disconnecting its internal light source and connecting an external mu…

Assessing Multilingual Fairness in Pre-trained Multimodal Representations

2021-06-12 · Findings (ACL) 2022 5 · Jialu Wang, Yang Liu, Xin Eric Wang

Recently pre-trained multimodal models, such as CLIP, have shown exceptional capabilities towards connecting images and natural language. The textual representations in English can be desirably transferred to multilingua…

Fairness

Connecting Vision and Language with Localized Narratives

2019-12-06 · ECCV 2020 8 · Jordi Pont-Tuset, Jasper Uijlings, Soravit Changpinyo, Radu Soricut 외

We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the regio…

FormImage CaptioningImage GenerationVisual Grounding