Data representation methods and use of mined corpora for Indian language transliteration
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalMachine TranslationTransliterationSimilar Papers 제목 키워드 기반
A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages
We take up the task of large-scale evaluation of neural machine transliteration between English and Indic languages, with a focus on multilingual transliteration to utilize orthographic similarity between Indian language…
TranslationTransliterationA Multilingual Parallel Corpora Collection Effort for Indian Languages
We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. Th…
Machine TranslationRetrievalSentenceTranslationCross-Corpora Language Recognition: A Preliminary Investigation with Indian Languages
In this paper, we conduct one of the very first studies for cross-corpora performance evaluation in the spoken language identification (LID) problem. Cross-corpora evaluation was not explored much in LID research, especi…
Language IdentificationSpoken language identificationAdvantages of Domain Knowledge Injection for Legal Document Summarization: A Case Study on Summarizing Indian Court Judgments in English and Hindi
Summarizing Indian legal court judgments is a complex task not only due to the intricate language and unstructured nature of the legal texts, but also since a large section of the Indian population does not understand th…
Document SummarizationAn Overview of Indian Spoken Language Recognition from Machine Learning Perspective
Automatic spoken language identification (LID) is a very important research field in the era of multilingual voice-command-based human-computer interaction (HCI). A front-end LID module helps to improve the performance o…
Language IdentificationSpoken language identification