paper-with-me

홈 › Papers

Improving Language Identification for Multilingual Speakers

2020-01-29 · Andrew Titus, Jan Silovsky, Nanxin Chen, Roger Hsiao, Mary Young, Arnab Ghoshal

Spoken language identification (LID) technologies have improved in recent years from discriminating largely distinct languages to discriminating highly similar languages or even dialects of the same language. One aspect that has been mostly neglected, however, is discrimination of languages for multilingual speakers, despite being a primary target audience of many systems that utilize LID technologies. As we show in this work, LID systems can have a high average accuracy for most combinations of languages while greatly underperforming for others when accented speech is present. We address this by using coarser-grained targets for the acoustic LID model and integrating its outputs with interaction context signals in a context-aware model to tailor the system to each user. This combined system achieves an average 97% accuracy across all language combinations while improving worst-case accuracy by over 60% relative to our baseline.

📄 PDF Abstract BibTeX arXiv:2001.11019

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationSpoken language identification

Similar Papers 제목 키워드 기반

Incorporating Dialectal Variability for Socially Equitable Language Identification

2017-07-01 · ACL 2017 7 · David Jurgens, Yulia Tsvetkov, Dan Jurafsky

Language identification (LID) is a critical first step for processing multilingual text. Yet most LID systems are not designed to handle the linguistic diversity of global platforms like Twitter, where local dialects and…

DiversityLanguage Identification

Identification of Languages in Algerian Arabic Multilingual Documents

2017-04-01 · WS 2017 4 · Wafia Adouane, Simon Dobnik

This paper presents a language identification system designed to detect the language of each word, in its context, in a multilingual documents as generated in social media by bilingual/multilingual communities, in our ca…

ChunkingGeneral ClassificationLanguage Identification

Whisper Speaker Identification: Leveraging Pre-Trained Multilingual Transformers for Robust Speaker Embeddings

2025-03-13 · Jakaria Islam Emon, Md Abu Salek, Kazi Tamanna Alam

Speaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data. In this paper, we propose WSI (Whisper Speaker Identification),…

Speaker Identificationspeech-recognitionSpeech Recognition

Multilingual and Cross-Lingual Complex Word Identification

2017-09-01 · RANLP 2017 9 · Seid Muhie Yimam, Sanja {\v{S}}tajner, Martin Riedl, Chris Biemann

Complex Word Identification (CWI) is an important task in lexical simplification and text accessibility. Due to the lack of CWI datasets, previous works largely depend on Simple English Wikipedia and edit histories for o…

Complex Word IdentificationLexical Simplification

GlobalPhone: Pronunciation Dictionaries in 20 Languages

2014-05-01 · LREC 2014 5 · Tanja Schultz, Tim Schlippe

This paper describes the advances in the multilingual text and speech database GlobalPhone, a multilingual database of high-quality read speech with corresponding transcriptions and pronunciation dictionaries in 20 langu…

Language IdentificationLanguage ModellingSpeaker Recognitionspeech-recognition+2