paper-with-me

Papers

Multi-Dialectal Representation Learning of Sinitic Phonology

2023-06-30 · Zhibai Jia

Machine learning techniques have shown their competence for representing and reasoning in symbolic systems such as language and phonology. In Sinitic Historical Phonology, notable tasks that could benefit from machine learning include the comparison of dialects and reconstruction of proto-languages systems. Motivated by this, this paper provides an approach for obtaining multi-dialectal representations of Sinitic syllables, by constructing a knowledge graph from structured phonological data, then applying the BoxE technique from knowledge base learning. We applied unsupervised clustering techniques to the obtained representations to observe that the representations capture phonemic contrast from the input dialects. Furthermore, we trained classifiers to perform inference of unobserved Middle Chinese labels, showing the representations' potential for indicating archaic, proto-language features. The representations can be used for performing completion of fragmented Sinitic phonological knowledge bases, estimating divergences between different characters, or aiding the exploration and reconstruction of archaic features.

📄 PDF Abstract BibTeX arXiv:2307.01209

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Conventional Orthography for Dialectal Arabic

2012-05-01 · LREC 2012 5 · Nizar Habash, Mona Diab, Owen Rambow

Dialectal Arabic (DA) refers to the day-to-day vernaculars spoken in the Arab world. DA lives side-by-side with the official language, Modern Standard Arabic (MSA). DA differs from MSA on all levels of linguistic represe…

Speech Recognition

Phonetic Modeling of Dialectal Variation in Vietnamese Speech

2026-05-23 · Quan Ngoc Hoang, Long Hoang Huu Nguyen, Nghia Hieu Nguyen, Kiet Van Nguyen 외 arxiv

Vietnamese exhibits substantial dialectal phonetic variation across Northern, Central, and Southern regions, where identical lexical items may be realized with markedly different pronunciations. Such variation poses chal…

Speech Recognition

North Sámi Dialect Identification with Self-supervised Speech Models

2023-05-19 · Sofoklis Kakouros, Katri Hiovain-Asikainen

The North S\'{a}mi (NS) language encapsulates four primary dialectal variants that are related but that also have differences in their phonology, morphology, and vocabulary. The unique geopolitical location of NS speaker…

Dialect Identification

A Unified Model for Arabizi Detection and Transliteration using Sequence-to-Sequence Models

2020-12-01 · COLING (WANLP) 2020 12 · Ali Shazal, Aiza Usman, Nizar Habash

While online Arabic is primarily written using the Arabic script, a Roman-script variety called Arabizi is often seen on social media. Although this representation captures the phonology of the language, it is not a one-…

Transliteration

Sinitic Wordnet: Laying the Groundwork with Chinese Varieties Written in Traditional Characters

2018-01-01 · GWC 2018 1 · Chih-Yao Lee, Shu-Kai Hsieh

The present work seeks to make the logographic nature of Chinese script a relevant research ground in wordnet studies. While wordnets are not so much about words as about the concepts represented in words, synset formati…