paper-with-me

홈 › Papers

Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine Translation

2023-05-23 · Di wu, Christof Monz

Using a vocabulary that is shared across languages is common practice in Multilingual Neural Machine Translation (MNMT). In addition to its simple design, shared tokens play an important role in positive knowledge transfer, assuming that shared tokens refer to similar meanings across languages. However, when word overlap is small, especially due to different writing systems, transfer is inhibited. In this paper, we define word-level information transfer pathways via word equivalence classes and rely on graph networks to fuse word embeddings across languages. Our experiments demonstrate the advantages of our approach: 1) embeddings of words with similar meanings are better aligned across languages, 2) our method achieves consistent BLEU improvements of up to 2.3 points for high- and low-resource MNMT, and 3) less than 1.0\% additional trainable parameters are required with a limited increase in computational costs, while inference time remains identical to the baseline. We release the codebase to the community.

📄 PDF Abstract BibTeX arXiv:2305.14189

Code (1)

moore3930/beyondsharedvocabulary 공식 구현 pytorch

Tasks

Machine TranslationTransfer LearningWord Embeddings

Similar Papers 제목 키워드 기반

Squeezing More from Limited Data with Recursive Transformers

2026-08-27 · Serdar Gülbahar, Lukas Edman, Alexander Fraser arxiv

Pre-training under limited data requires a different view of scaling than web-scale language modeling. With a fixed data budget but relatively abundant compute, increasing parameter count helps only up to an optimal scal…

From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs

2026-01-31 · Louis Schiekiera, Max Zimmer, Christophe Roux, Sebastian Pokutta 외 arxiv

We investigate the extent to which an LLM's hidden-state geometry can be recovered from its behavior in psycholinguistic experiments. Across eight instruction-tuned transformer models, we run two experimental paradigms -…

Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI

2026-07-10 · Yuan Cao, Haiqian Yang arxiv

Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: t…

LightRNN: Memory and Computation-Efficient Recurrent Neural Networks

2016-10-31 · NeurIPS 2016 12 · Xiang Li, Tao Qin, Jian Yang, Tie-Yan Liu

Recurrent neural networks (RNNs) have achieved state-of-the-art performances in many natural language processing tasks, such as language modeling and machine translation. However, when the vocabulary is large, the RNN mo…

GPULanguage ModelingLanguage ModellingMachine Translation

Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models

2025-09-25 · Zoe Wanying He, Sean Trott, Meenakshi Khosla arxiv

Recent studies show that deep vision-only and language-only models--trained on disjoint modalities--nonetheless project their inputs into a partially aligned representational space. Yet we still lack a clear picture of w…