paper-with-me

홈 › Papers

Identification/Segmentation of Indian Regional Languages with Singular Value Decomposition based Feature Embedding

2020-05-17

language identification (LID) is identifing a language in a given spoken utterance. Language segmentation is equally inportant as language identification where language boundaries can be spotted in a multi language utterance. In this paper, we have experimented with two schemes for language identification in Indian regional language context as very few works has been done. Singular value based feature embedding is used for both of the schemes. In first scheme, the singular value decomposition (SVD) is applied to the n-gram utterance matrix and in the second scheme, SVD is applied on the difference supervector matrix space. We have observed that in both the schemes, 55-65% singular value energy is sufficient to capture the language context. In n-gram based feature representation, we have seen that different skipgram models capture different language context. We have observed that for short test duration, supervector based feature representation is better but with a longer duration test signal, n-gram based feature performed better. We have also extended our work to explore language-based segmentation where we have seen that segmentation accuracy of four language group with ten language training model, scheme-1 has performed well but with same four language training model, scheme-2 outperformed scheme-1

📄 PDF Abstract BibTeX arXiv:2005.08229

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationSegmentation

Similar Papers 제목 키워드 기반

Indian Regional Movie Dataset for Recommender Systems

2018-01-07 · Prerna Agarwal, Richa Verma, Angshul Majumdar

Indian regional movie dataset is the first database of regional Indian movies, users and their ratings. It consists of movies belonging to 18 different Indian regional languages and metadata of users with varying demogra…

Collaborative Filteringcompressed sensingDiversityMatrix Completion+1

Non-native Accent Partitioning for Speakers of Indian Regional Languages

2019-12-01 · ICON 2019 12 · Radha Krishna Guntur, Krishnan Ramakrishnan, Vinay Kumar Mittal

Acoustic features extracted from the speech signal can help in identifying speaker related multiple information such as geographical origin, regional accent and nativity. In this paper, classification of native speakers …

Labeling of Query Words using Conditional Random Field

2016-07-29 · Satanu Ghosh, Souvick Ghosh, Dipankar Das

This paper describes our approach on Query Word Labeling as an attempt in the shared task on Mixed Script Information Retrieval at Forum for Information Retrieval Evaluation (FIRE) 2015. The query is written in Roman scr…

Information RetrievalLanguage IdentificationRetrieval

Survey of Pseudonymization, Abstractive Summarization & Spell Checker for Hindi and Marathi

2024-12-24 · Rasika Ransing, Mohammed Amaan Dhamaskar, Ayush Rajpurohit, Amey Dhoke 외

India's vast linguistic diversity presents unique challenges and opportunities for technological advancement, especially in the realm of Natural Language Processing (NLP). While there has been significant progress in NLP…

Abstractive Text SummarizationDiversityText AnonymizationText Summarization

muBoost: An Effective Method for Solving Indic Multilingual Text Classification Problem

2022-06-21 · Manish Pathak, Aditya Jain

Text Classification is an integral part of many Natural Language Processing tasks such as sarcasm detection, sentiment analysis and many more such applications. Many e-commerce websites, social-media/entertainment platfo…

Multilingual text classificationSarcasm DetectionSentiment Analysistext-classification+1