paper-with-me

홈 › Papers

Jambu: A historical linguistic database for South Asian languages

2023-06-05 · Aryaman Arora, Adam Farris, Samopriya Basu, Suresh Kolichala

We introduce Jambu, a cognate database of South Asian languages which unifies dozens of previous sources in a structured and accessible format. The database includes 287k lemmata from 602 lects, grouped together in 23k sets of cognates. We outline the data wrangling necessary to compile the dataset and train neural models for reflex prediction on the Indo-Aryan subset of the data. We hope that Jambu is an invaluable resource for all historical linguists and Indologists, and look towards further improvement and expansion of the database.

📄 PDF Abstract BibTeX arXiv:2306.02514

Code (1)

moli-mandala/data 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Computational historical linguistics and language diversity in South Asia

2022-03-23 · ACL 2022 5 · Aryaman Arora, Adam Farris, Samopriya Basu, Suresh Kolichala

South Asia is home to a plethora of languages, many of which severely lack access to new language technologies. This linguistic diversity also results in a research environment conducive to the study of comparative, cont…

Diversity

Computational historical linguistics and language diversity in South Asia

2021-11-16 · ACL ARR November 2021 11 · Anonymous

South Asia is home to a plethora of languages, most of which are severely lacking access to language technologies that have been developed with the maturity of NLP/CL. This linguistic diversity, however, also results in …

Diversity

Bhasacitra: Visualising the dialect geography of South Asia

2021-05-28 · Aryaman Arora, Adam Farris, Gopalakrishnan R, Samopriya Basu

We present Bhasacitra, a dialect mapping system for South Asia built on a database of linguistic studies of languages of the region annotated for topic and location data. We analyse language coverage and look towards app…

Bhāṣācitra: Visualising the dialect geography of South Asia

2021-08-01 · ACL (LChange) 2021 8 · Aryaman Arora, Adam Farris, Gopalakrishnan R, Samopriya Basu

We present Bhāṣācitra, a dialect mapping system for South Asia built on a database of linguistic studies of languages of the region annotated for topic and location data. We analyse language coverage and look towards app…

Cloud-based Automatic Speech Recognition Systems for Southeast Asian Languages

2022-10-07 · Lei Wang, Rong Tong, Cheung Chi Leung, Sunil Sivadas 외

This paper provides an overall introduction of our Automatic Speech Recognition (ASR) systems for Southeast Asian languages. As not much existing work has been carried out on such regional languages, a few difficulties s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition