paper-with-me

홈 › Papers

Should All Cross-Lingual Embeddings Speak English?

2019-11-08 · ACL 2020 6 · Antonios Anastasopoulos, Graham Neubig

Most of recent work in cross-lingual word embeddings is severely Anglocentric. The vast majority of lexicon induction evaluation dictionaries are between English and another language, and the English embedding space is selected by default as the hub when learning in a multilingual setting. With this work, however, we challenge these practices. First, we show that the choice of hub language can significantly impact downstream lexicon induction performance. Second, we both expand the current evaluation dictionary collection to include all language pairs using triangulation, and also create new dictionaries for under-represented languages. Evaluating established methods over all these language pairs sheds light into their suitability and presents new challenges for the field. Finally, in our analysis we identify general guidelines for strong cross-lingual embeddings baselines, based on more than just Anglocentric experiments.

📄 PDF Abstract BibTeX arXiv:1911.03058

Code (1)

antonisa/embeddings 공식 구현

Tasks

AllCross-Lingual Word EmbeddingsWord Embeddings

Similar Papers 제목 키워드 기반

Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings

2025-03-13 · Jakaria Islam Emon, Md Abu Salek, Kazi Tamanna Alam

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification),…

Speaker Identificationspeech-recognitionSpeech Recognition

Cross-lingual Multispeaker Text-to-Speech under Limited-Data Scenario

2020-05-21 · Zexin Cai, Yaogen Yang, Ming Li

Modeling voices for multiple speakers and multiple languages in one text-to-speech system has been a challenge for a long time. This paper presents an extension on Tacotron2 to achieve bilingual multispeaker speech synth…

AttributeSpeech Synthesistext-to-speechText to Speech

Sentiment Analysis for Hinglish Code-mixed Tweets by means of Cross-lingual Word Embeddings

2020-05-01 · LREC 2020 5 · Pranaydeep Singh, Els Lefever

This paper investigates the use of unsupervised cross-lingual embeddings for solving the problem of code-mixed social media text understanding. We specifically investigate the use of these embeddings for a sentiment anal…

Cross-Lingual Word EmbeddingsSentiment AnalysisTransfer LearningWord Embeddings

First Bilingual Word Embeddings for te reo Māori and English: Towards Code-switching Detection in a Low-resourced setting

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Māori speakers are bilingual, where Māori is code-switched with English. With Māori being low-resourced for technology development, there are minimal resources available for Māori-English code-switch detection. This res…

Word Embeddings

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

2026-07-03 · Anisha Pattanayak, Hanie Kang, Huang-Cheng Chou, Shrikanth Narayanan 외 hf

Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic populations. Speech-based depression detection performs well monolingually, but cross-lingual generalizat…