paper-with-me

Papers Spoken language identification

“Spoken language identification” 태그가 달린 논문 53편 · 필터 해제

Spoken Language Identification with Pre-trained Models and Margin Loss

2026-05-03 · Zhihua Fang, Liang He, Weiwu Jiang arxiv

For the speaker-controlled spoken language identification task proposed in the TidyLang Challenge 2026, this paper proposes a language identification method based on pre-trained models and margin-based losses. The propos…

Spoken language identification

Geolocation-Aware Robust Spoken Language Identification

2025-08-23 · Qingzheng Wang, Hye-jin Shim, Jiancheng Sun, Shinji Watanabe arxiv

While Self-supervised Learning (SSL) has significantly improved Spoken Language Identification (LID), existing models often struggle to consistently classify dialects and accents of the same language as a unified class. …

Spoken language identificationSelf-Supervised Learning

On the use of Performer and Agent Attention for Spoken Language Identification

2025-02-09 · Jitendra Kumar dhiman, Jainag Ambati

One of the methods for language Identification (LID) involves deriving speech representation from pre-trained models using self-supervised learning, followed by fine-tuning the model for the LID task. State-of-the-art ap…

Language IdentificationSelf-Supervised LearningSpoken language identification

AfriHuBERT: A self-supervised speech representation model for African languages

2024-09-30 · Jesujoba O. Alabi, Xuechen Liu, Dietrich Klakow, Junichi Yamagishi

In this work, we present AfriHuBERT, an extension of mHuBERT-147, a compact self-supervised learning (SSL) model pretrained on 147 languages. While mHuBERT-147 covered 16 African languages, we expand this to 1,226 throug…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-corpusLanguage Identification+4

Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking

2024-09-27 · Brian Yan, Vineel Pratap, Shinji Watanabe, Michael Auli

Multilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often not the case for most practical setting…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationRe-Ranking+3

Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech

2024-06-13 · Martina Valente, Fabio Brugnara, Giovanni Morrone, Enrico Zovato 외

This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Obser…

Language Identificationspeaker-diarizationSpeaker Diarizationspeech-recognition+2

Generative linguistic representation for spoken language identification

2023-12-18 · Peng Shen, Xuguang Lu, Hisashi Kawai

Effective extraction and application of linguistic features are central to the enhancement of spoken Language IDentification (LID) performance. With the success of recent large models, such as GPT and Whisper, the potent…

DecoderLanguage Identificationspeech-recognitionSpeech Recognition+1

Self-supervised Adaptive Pre-training of Multilingual Speech Models for Language and Dialect Identification

2023-12-12 · Mohammed Maqsood Shaik, Dietrich Klakow, Badr M. Abdullah

Pre-trained Transformer-based speech models have shown striking performance when fine-tuned on various downstream tasks such as automatic speech recognition and spoken language identification (SLID). However, the problem…

Automatic Speech RecognitionDialect IdentificationFew-Shot LearningLanguage Identification+3

Wavelet Scattering Transform for Improving Generalization in Low-Resourced Spoken Language Identification

2023-10-01 · Spandan Dey, Premjeet Singh, Goutam Saha

Commonly used features in spoken language identification (LID), such as mel-spectrogram or MFCC, lose high-frequency information due to windowing. The loss further increases for longer temporal contexts. To improve gener…

Language IdentificationSpoken language identification

Multimodal Modeling For Spoken Language Identification

2023-09-19 · Shikhar Bharadwaj, Min Ma, Shikhar Vashishth, Ankur Bapna 외

Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have …

Language IdentificationSpoken language identification

Robust Open-Set Spoken Language Identification and the CU MultiLang Dataset

2023-08-29 · Mustafa Eyceoz, Justin Lee, Siddharth Pittie, Homayoon Beigi

Most state-of-the-art spoken language identification models are closed-set; in other words, they can only output a language label from the set of classes they were trained on. Open-set spoken language identification syst…

Language IdentificationSpoken language identification

Unified model for code-switching speech recognition and language identification based on a concatenated tokenizer

2023-06-14 · Kunal Dhawan, Dima Rekesh, Boris Ginsburg

Code-Switching (CS) multilingual Automatic Speech Recognition (ASR) models can transcribe speech containing two or more alternating languages during a conversation. This paper proposes (1) a new method for creating code-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+2

Spoken Language Identification System for English-Mandarin Code-Switching Child-Directed Speech

2023-06-01 · Shashi Kant Gupta, Sushant Hiray, Prashant Kukde

This work focuses on improving the Spoken Language Identification (LangId) system for a challenge that focuses on developing robust language identification systems that are reliable for non-standard, accented (Singaporea…

DecoderLanguage IdentificationSpoken language identification

Improving Spoken Language Identification with Map-Mix

2023-02-16 · Shangeth Rajaa, Kriti Anandan, Swaraj Dalmia, Tarun Gupta 외

The pre-trained multi-lingual XLSR model generalizes well for language identification after fine-tuning on unseen languages. However, the performance significantly degrades when the languages are not very distinct from e…

Data AugmentationLanguage IdentificationSpoken language identification

Cross-Corpora Spoken Language Identification with Domain Diversification and Generalization

2023-02-10 · Spandan Dey, Md Sahidullah, Goutam Saha

This work addresses the cross-corpora generalization issue for the low-resourced spoken language identification (LID) problem. We have conducted the experiments in the context of Indian LID and identified strikingly poor…

Data AugmentationDomain GeneralizationLanguage IdentificationSpoken language identification

An Overview of Indian Spoken Language Recognition from Machine Learning Perspective

2022-11-30 · Spandan Dey, Md Sahidullah, Goutam Saha

Automatic spoken language identification (LID) is a very important research field in the era of multilingual voice-command-based human-computer interaction (HCI). A front-end LID module helps to improve the performance o…

Language IdentificationSpoken language identification

Accidental Learners: Spoken Language Identification in Multilingual Self-Supervised Models

2022-11-09 · Travis M. Bartley, Fei Jia, Krishna C. Puvvada, Samuel Kriman 외

In this paper, we extend previous self-supervised approaches for language identification by experimenting with Conformer based architecture in a multilingual pre-training paradigm. We find that pre-trained speech models …

Language IdentificationSpoken language identification

A Compact End-to-End Model with Local and Global Context for Spoken Language Identification

2022-10-27 · Fei Jia, Nithin Rao Koluguri, Jagadeesh Balam, Boris Ginsburg

We introduce TitaNet-LID, a compact end-to-end neural network for Spoken Language Identification (LID) that is based on the ContextNet architecture. TitaNet-LID employs 1D depth-wise separable convolutions and Squeeze-an…

Language IdentificationSpoken language identification

EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use

2022-07-12 · Jan Schlüter, Gerald Gutenbrunner

In audio classification, differentiable auditory filterbanks with few parameters cover the middle ground between hard-coded spectrograms and raw audio. LEAF (arXiv:2101.08596), a Gabor-based filterbank combined with Per-…

Audio ClassificationClassificationInstrument RecognitionPitch Classification+1

Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource Devices

2022-07-12 · Harlin Lee, Aaqib Saeed

This work introduces BRILLsson, a novel binary neural network-based representation learning model for a broad range of non-semantic speech tasks. We train the model with knowledge distillation from a large and real-value…

Emotion RecognitionKeyword SpottingKnowledge DistillationLanguage Identification+2
1–20 / 53 다음 →