Bridging Languages through Images with Deep Partial Canonical Correlation Analysis
We present a deep neural network that leverages images to improve bilingual text embeddings. Relying on bilingual image tags and descriptions, our approach conditions text embedding induction on the shared visual information for both languages, producing highly correlated bilingual embeddings. In particular, we propose a novel model based on Partial Canonical Correlation Analysis (PCCA). While the original PCCA finds linear projections of two views in order to maximize their canonical correlation conditioned on a shared third variable, we introduce a non-linear Deep PCCA (DPCCA) model, and develop a new stochastic iterative algorithm for its optimization. We evaluate PCCA and DPCCA on multilingual word similarity and cross-lingual image description retrieval. Our models outperform a large variety of previous methods, despite not having access to any visual signal during test time inference.
Code (1)
Tasks
Image DescriptionImage RetrievalQuestion AnsweringRepresentation LearningRetrievalVisual Question Answering (VQA)Word SimilaritySimilar Papers 제목 키워드 기반
ConDor: Self-Supervised Canonicalization of 3D Pose for Partial Shapes
Progress in 3D object understanding has relied on manually canonicalized shape datasets that contain instances with consistent position and orientation (3D pose). This has made it hard to generalize these methods to in-t…
3D Canonicalization3D Geometry Perception3D Part Segmentation3D Pose Estimation+1Why are some word orders more common than others? A uniform information density account
Languages vary widely in many ways, including their canonical word order. A basic aspect of the observed variation is the fact that some word orders are much more common than others. Although this regularity has been rec…
SentenceBridgeDiff: Bridging Human Observations and Flat-Garment Synthesis for Virtual Try-Off
Virtual try-off (VTOFF) aims to recover canonical flat-garment representations from images of dressed persons for standardized display and downstream virtual try-on. Prior methods often treat VTOFF as direct image transl…
Virtual Try-OffVirtual Try-onBridging Languages through Etymology: The case of cross language text categorization
Weakly-supervised 3D Shape Completion in the Wild
3D shape completion for real data is important but challenging, since partial point clouds acquired by real-world sensors are usually sparse, noisy and unaligned. Different from previous methods, we address the problem o…
Point Cloud RegistrationPose Estimation