paper-with-me

Papers

Properties of phoneme N -grams across the world's language families

2014-01-04 · Taraka Rama, Lars Borin

In this article, we investigate the properties of phoneme N-grams across half of the world's languages. We investigate if the sizes of three different N-gram distributions of the world's language families obey a power law. Further, the N-gram distributions of language families parallel the sizes of the families, which seem to obey a power law distribution. The correlation between N-gram distributions and language family sizes improves with increasing values of N. We applied statistical tests, originally given by physicists, to test the hypothesis of power law fit to twelve different datasets. The study also raises some new questions about the use of N-gram distributions in linguistic research, which we answer by running a statistical test.

📄 PDF Abstract BibTeX arXiv:1401.0794

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation

2025-04-11 · Haowei Lou, Hye-Young Paik, Sheng Li, Wen Hu 외

Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging due to discrepancies in phoneme vocabul…

text-to-speechText to Speech

German Phoneme Recognition with Text-to-Phoneme Data Augmentation

2022-11-24 · Dojun Park, Seohyun Park

In this study, we experimented to examine the effect of adding the most frequent n phoneme bigrams to the basic vocabulary on the German phoneme recognition model using the text-to-phoneme data augmentation strategy. As …

Data AugmentationPhoneme Recognition

The Distribution of Phoneme Frequencies across the World's Languages: Macroscopic and Microscopic Information-Theoretic Models

2026-03-03 · Fermín Moscoso del Prado Martín, Suchir Salhan arxiv

We demonstrate that the frequency distribution of phonemes across languages can be explained at both macroscopic and microscopic levels. Macroscopically, phoneme rank-frequency distributions closely follow the order stat…

PHO-LID: A Unified Model Incorporating Acoustic-Phonetic and Phonotactic Information for Language Identification

2022-03-23 · Hexin Liu, Leibny Paola Garcia Perera, Andy W. H. Khong, Suzy J. Styles 외

We propose a novel model to hierarchically incorporate phoneme and phonotactic information for language identification (LID) without requiring phoneme annotations for training. In this model, named PHO-LID, a self-superv…

Language Identification

IPA-CHILDES & G2P+: Feature-Rich Resources for Cross-Lingual Phonology and Phonemic Language Modeling

2025-04-03 · Zébulon Goriely, Paula Buttery

In this paper, we introduce two resources: (i) G2P+, a tool for converting orthographic datasets to a consistent phonemic representation; and (ii) IPA CHILDES, a phonemic dataset of child-centered speech across 31 langua…

Grapheme-to-Phoneme ConversionLanguage ModelingLanguage Modelling