Papers Spoken language identification
“Spoken language identification” 태그가 달린 논문 53편 · 필터 해제
Spoken Language Identification with Pre-trained Models and Margin Loss
For the speaker-controlled spoken language identification task proposed in the TidyLang Challenge 2026, this paper proposes a language identification method based on pre-trained models and margin-based losses. The propos…
Spoken language identificationGeolocation-Aware Robust Spoken Language Identification
While Self-supervised Learning (SSL) has significantly improved Spoken Language Identification (LID), existing models often struggle to consistently classify dialects and accents of the same language as a unified class. …
Spoken language identificationSelf-Supervised LearningOn the use of Performer and Agent Attention for Spoken Language Identification
One of the methods for language Identification (LID) involves deriving speech representation from pre-trained models using self-supervised learning, followed by fine-tuning the model for the LID task. State-of-the-art ap…
Language IdentificationSelf-Supervised LearningSpoken language identificationAfriHuBERT: A self-supervised speech representation model for African languages
In this work, we present AfriHuBERT, an extension of mHuBERT-147, a compact self-supervised learning (SSL) model pretrained on 147 languages. While mHuBERT-147 covered 16 African languages, we expand this to 1,226 throug…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-corpusLanguage Identification+4Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
Multilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often not the case for most practical setting…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationRe-Ranking+3Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Obser…
Language Identificationspeaker-diarizationSpeaker Diarizationspeech-recognition+2Generative linguistic representation for spoken language identification
Effective extraction and application of linguistic features are central to the enhancement of spoken Language IDentification (LID) performance. With the success of recent large models, such as GPT and Whisper, the potent…
DecoderLanguage Identificationspeech-recognitionSpeech Recognition+1Self-supervised Adaptive Pre-training of Multilingual Speech Models for Language and Dialect Identification
Pre-trained Transformer-based speech models have shown striking performance when fine-tuned on various downstream tasks such as automatic speech recognition and spoken language identification (SLID). However, the problem…
Automatic Speech RecognitionDialect IdentificationFew-Shot LearningLanguage Identification+3Wavelet Scattering Transform for Improving Generalization in Low-Resourced Spoken Language Identification
Commonly used features in spoken language identification (LID), such as mel-spectrogram or MFCC, lose high-frequency information due to windowing. The loss further increases for longer temporal contexts. To improve gener…
Language IdentificationSpoken language identificationMultimodal Modeling For Spoken Language Identification
Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have …
Language IdentificationSpoken language identificationRobust Open-Set Spoken Language Identification and the CU MultiLang Dataset
Most state-of-the-art spoken language identification models are closed-set; in other words, they can only output a language label from the set of classes they were trained on. Open-set spoken language identification syst…
Language IdentificationSpoken language identificationUnified model for code-switching speech recognition and language identification based on a concatenated tokenizer
Code-Switching (CS) multilingual Automatic Speech Recognition (ASR) models can transcribe speech containing two or more alternating languages during a conversation. This paper proposes (1) a new method for creating code-…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+2Spoken Language Identification System for English-Mandarin Code-Switching Child-Directed Speech
This work focuses on improving the Spoken Language Identification (LangId) system for a challenge that focuses on developing robust language identification systems that are reliable for non-standard, accented (Singaporea…
DecoderLanguage IdentificationSpoken language identificationImproving Spoken Language Identification with Map-Mix
The pre-trained multi-lingual XLSR model generalizes well for language identification after fine-tuning on unseen languages. However, the performance significantly degrades when the languages are not very distinct from e…
Data AugmentationLanguage IdentificationSpoken language identificationCross-Corpora Spoken Language Identification with Domain Diversification and Generalization
This work addresses the cross-corpora generalization issue for the low-resourced spoken language identification (LID) problem. We have conducted the experiments in the context of Indian LID and identified strikingly poor…
Data AugmentationDomain GeneralizationLanguage IdentificationSpoken language identificationAn Overview of Indian Spoken Language Recognition from Machine Learning Perspective
Automatic spoken language identification (LID) is a very important research field in the era of multilingual voice-command-based human-computer interaction (HCI). A front-end LID module helps to improve the performance o…
Language IdentificationSpoken language identificationAccidental Learners: Spoken Language Identification in Multilingual Self-Supervised Models
In this paper, we extend previous self-supervised approaches for language identification by experimenting with Conformer based architecture in a multilingual pre-training paradigm. We find that pre-trained speech models …
Language IdentificationSpoken language identificationA Compact End-to-End Model with Local and Global Context for Spoken Language Identification
We introduce TitaNet-LID, a compact end-to-end neural network for Spoken Language Identification (LID) that is based on the ContextNet architecture. TitaNet-LID employs 1D depth-wise separable convolutions and Squeeze-an…
Language IdentificationSpoken language identificationEfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use
In audio classification, differentiable auditory filterbanks with few parameters cover the middle ground between hard-coded spectrograms and raw audio. LEAF (arXiv:2101.08596), a Gabor-based filterbank combined with Per-…
Audio ClassificationClassificationInstrument RecognitionPitch Classification+1Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource Devices
This work introduces BRILLsson, a novel binary neural network-based representation learning model for a broad range of non-semantic speech tasks. We train the model with knowledge distillation from a large and real-value…
Emotion RecognitionKeyword SpottingKnowledge DistillationLanguage Identification+2