paper-with-me

홈 › Papers

Phonotactic Complexity across Dialects

2024-02-20 · Ryan Soh-Eun Shim, Kalvin Chang, David R. Mortensen

Received wisdom in linguistic typology holds that if the structure of a language becomes more complex in one dimension, it will simplify in another, building on the assumption that all languages are equally complex (Joseph and Newmeyer, 2012). We study this claim on a micro-level, using a tightly-controlled sample of Dutch dialects (across 366 collection sites) and Min dialects (across 60 sites), which enables a more fair comparison across varieties. Even at the dialect level, we find empirical evidence for a tradeoff between word length and a computational measure of phonotactic complexity from a LSTM-based phone-level language model-a result previously documented only at the language level. A generalized additive model (GAM) shows that dialects with low phonotactic complexity concentrate around the capital regions, which we hypothesize to correspond to prior hypotheses that language varieties of greater or more diverse populations show reduced phonotactic complexity. We also experiment with incorporating the auxiliary task of predicting syllable constituency, but do not find an increase in the negative correlation observed.

📄 PDF Abstract BibTeX arXiv:2402.12998

Code (1)

cmu-llab/phonotactic-complexity-across-dialects 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Rethinking Phonotactic Complexity

2019-08-01 · WS 2019 8 · Tiago Pimentel, Brian Roark, Ryan Cotterell

In this work, we propose the use of phone-level language models to estimate phonotactic complexity{---}measured in bits per phoneme{---}which makes cross-linguistic comparison straightforward. We compare the entropy acro…

Multi-view Dimensionality Reduction for Dialect Identification of Arabic Broadcast Speech

2016-09-19 · Sameer Khurana, Ahmed Ali, Steve Renals

In this work, we present a new Vector Space Model (VSM) of speech utterances for the task of spoken dialect identification. Generally, DID systems are built using two sets of features that are extracted from speech utter…

Dialect IdentificationDimensionality Reduction

Phonotactic Complexity and its Trade-offs

2020-05-07 · TACL 2020 1 · Tiago Pimentel, Brian Roark, Ryan Cotterell

We present methods for calculating a measure of phonotactic complexity---bits per phoneme---that permits a straightforward cross-linguistic comparison. When given a word, represented as a sequence of phonemic segments su…

Literary and Colloquial Dialect Identification for Tamil using Acoustic Features

2024-08-27 · M. Nanmalar, P. Vijayalakshmi, T. Nagarajan

The evolution and diversity of a language is evident from it's various dialects. If the various dialects are not addressed in technological advancements like automatic speech recognition and speech synthesis, there is a …

Automatic Speech RecognitionDialect IdentificationLanguage Identificationspeech-recognition+2

Correlation Does Not Imply Compensation: Complexity and Irregularity in the Lexicon

2024-06-07 · Amanda Doucette, Ryan Cotterell, Morgan Sonderegger, Timothy J. O'Donnell

It has been claimed that within a language, morphologically irregular words are more likely to be phonotactically simple and morphologically regular words are more likely to be phonotactically complex. This inverse corre…