paper-with-me

홈 › Papers

Disentangling dialects: a neural approach to Indo-Aryan historical phonology and subgrouping

2020-11-01 · CONLL 2020 · Chundra Cathcart, Taraka Rama

This paper seeks to uncover patterns of sound change across Indo-Aryan languages using an LSTM encoder-decoder architecture. We augment our models with embeddings represent-ing language ID, part of speech, and other features such as word embeddings. We find that a highly augmented model shows highest accuracy in predicting held-out forms, and investigate other properties of interest learned by our models{'} representations. We outline extensions to this architecture that can better capture variation in Indo-Aryan sound change.

📄 PDF Abstract BibTeX

Code (1)

chundrac/ia-conll-2020 공식 구현 tf

Tasks

DecoderWord Embeddings

Methods 이 논문이 사용한 방법론

Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Toward a deep dialectological representation of Indo-Aryan

2019-06-01 · WS 2019 6 · Chundra Cathcart

This paper presents a new approach to disentangling inter-dialectal and intra-dialectal relationships within one such group, the Indo-Aryan subgroup of Indo-European. We draw upon admixture models and deep generative mod…

Multi-Dialectal Representation Learning of Sinitic Phonology

2023-06-30 · Zhibai Jia

Machine learning techniques have shown their competence for representing and reasoning in symbolic systems such as language and phonology. In Sinitic Historical Phonology, notable tasks that could benefit from machine le…

Representation Learning

Detection of Similar Languages and Dialects Using Deep Supervised Autoencoder

2020-12-01 · ICON 2020 12 · Shantipriya Parida, Esau Villatoro-Tello, Sajit Kumar, Maël Fabien 외

Language detection is considered a difficult task especially for similar languages, varieties, and dialects. With the growing number of online content in different languages, the need for reliable and robust language det…

Language Identification

Jambu: A historical linguistic database for South Asian languages

2023-06-05 · Aryaman Arora, Adam Farris, Samopriya Basu, Suresh Kolichala

We introduce Jambu, a cognate database of South Asian languages which unifies dozens of previous sources in a structured and accessible format. The database includes 287k lemmata from 602 lects, grouped together in 23k s…

IndoNLP 2025: Shared Task on Real-Time Reverse Transliteration for Romanized Indo-Aryan languages

2025-01-10 · Deshan Sumanathilaka, Isuri Anuradha, Ruvan Weerasinghe, Nicholas Micallef 외

The paper overviews the shared task on Real-Time Reverse Transliteration for Romanized Indo-Aryan languages. It focuses on the reverse transliteration of low-resourced languages in the Indo-Aryan family to their native s…

Transliteration