Identification/Segmentation of Indian Regional Languages with Singular Value Decomposition based Feature Embedding
language identification (LID) is identifing a language in a given spoken utterance. Language segmentation is equally inportant as language identification where language boundaries can be spotted in a multi language utterance. In this paper, we have experimented with two schemes for language identification in Indian regional language context as very few works has been done. Singular value based feature embedding is used for both of the schemes. In first scheme, the singular value decomposition (SVD) is applied to the n-gram utterance matrix and in the second scheme, SVD is applied on the difference supervector matrix space. We have observed that in both the schemes, 55-65% singular value energy is sufficient to capture the language context. In n-gram based feature representation, we have seen that different skipgram models capture different language context. We have observed that for short test duration, supervector based feature representation is better but with a longer duration test signal, n-gram based feature performed better. We have also extended our work to explore language-based segmentation where we have seen that segmentation accuracy of four language group with ten language training model, scheme-1 has performed well but with same four language training model, scheme-2 outperformed scheme-1
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationSegmentationSimilar Papers 제목 키워드 기반
Indian Regional Movie Dataset for Recommender Systems
Indian regional movie dataset is the first database of regional Indian movies, users and their ratings. It consists of movies belonging to 18 different Indian regional languages and metadata of users with varying demogra…
Collaborative Filteringcompressed sensingDiversityMatrix Completion+1Non-native Accent Partitioning for Speakers of Indian Regional Languages
Acoustic features extracted from the speech signal can help in identifying speaker related multiple information such as geographical origin, regional accent and nativity. In this paper, classification of native speakers …
Labeling of Query Words using Conditional Random Field
This paper describes our approach on Query Word Labeling as an attempt in the shared task on Mixed Script Information Retrieval at Forum for Information Retrieval Evaluation (FIRE) 2015. The query is written in Roman scr…
Information RetrievalLanguage IdentificationRetrievalSurvey of Pseudonymization, Abstractive Summarization & Spell Checker for Hindi and Marathi
India's vast linguistic diversity presents unique challenges and opportunities for technological advancement, especially in the realm of Natural Language Processing (NLP). While there has been significant progress in NLP…
Abstractive Text SummarizationDiversityText AnonymizationText SummarizationmuBoost: An Effective Method for Solving Indic Multilingual Text Classification Problem
Text Classification is an integral part of many Natural Language Processing tasks such as sarcasm detection, sentiment analysis and many more such applications. Many e-commerce websites, social-media/entertainment platfo…
Multilingual text classificationSarcasm DetectionSentiment Analysistext-classification+1