Should All Cross-Lingual Embeddings Speak English?
Most of recent work in cross-lingual word embeddings is severely Anglocentric. The vast majority of lexicon induction evaluation dictionaries are between English and another language, and the English embedding space is selected by default as the hub when learning in a multilingual setting. With this work, however, we challenge these practices. First, we show that the choice of hub language can significantly impact downstream lexicon induction performance. Second, we both expand the current evaluation dictionary collection to include all language pairs using triangulation, and also create new dictionaries for under-represented languages. Evaluating established methods over all these language pairs sheds light into their suitability and presents new challenges for the field. Finally, in our analysis we identify general guidelines for strong cross-lingual embeddings baselines, based on more than just Anglocentric experiments.
Code (1)
Tasks
AllCross-Lingual Word EmbeddingsWord EmbeddingsSimilar Papers 제목 키워드 기반
Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings
Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification),…
Speaker Identificationspeech-recognitionSpeech RecognitionCross-lingual Multispeaker Text-to-Speech under Limited-Data Scenario
Modeling voices for multiple speakers and multiple languages in one text-to-speech system has been a challenge for a long time. This paper presents an extension on Tacotron2 to achieve bilingual multispeaker speech synth…
AttributeSpeech Synthesistext-to-speechText to SpeechSentiment Analysis for Hinglish Code-mixed Tweets by means of Cross-lingual Word Embeddings
This paper investigates the use of unsupervised cross-lingual embeddings for solving the problem of code-mixed social media text understanding. We specifically investigate the use of these embeddings for a sentiment anal…
Cross-Lingual Word EmbeddingsSentiment AnalysisTransfer LearningWord EmbeddingsFirst Bilingual Word Embeddings for te reo Māori and English: Towards Code-switching Detection in a Low-resourced setting
Māori speakers are bilingual, where Māori is code-switched with English. With Māori being low-resourced for technology development, there are minimal resources available for Māori-English code-switch detection. This res…
Word EmbeddingsLayer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment
Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic populations. Speech-based depression detection performs well monolingually, but cross-lingual generalizat…