paper-with-me

Papers

SegBo: A Database of Borrowed Sounds in the World's Languages

2020-05-01 · LREC 2020 5 · Eitan Grossman, Elad Eisen, Dmitry Nikolaev, Steven Moran

Phonological segment borrowing is a process through which languages acquire new contrastive speech sounds as the result of borrowing new words from other languages. Despite the fact that phonological segment borrowing is documented in many of the world{'}s languages, to date there has been no large-scale quantitative study of the phenomenon. In this paper, we present SegBo, a novel cross-linguistic database of borrowed phonological segments. We describe our data aggregation pipeline and the resulting language sample. We also present two short case studies based on the database. The first deals with the impact of large colonial languages on the sound systems of the world{'}s languages; the second deals with universals of borrowing in the domain of rhotic consonants.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic

2024-08-05 · Yassine El Kheir, Hamdy Mubarak, Ahmed Ali, Shammur Absar Chowdhury

This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its stand…

Dialect2SQL: A Novel Text-to-SQL Dataset for Arabic Dialects with a Focus on Moroccan Darija

2025-01-20 · Salmane Chafik, Saad Ezzini, Ismail Berrada

The task of converting natural language questions (NLQs) into executable SQL queries, known as text-to-SQL, has gained significant interest in recent years, as it enables non-technical users to interact with relational d…

Text to SQLText-To-SQL

A study for the effect of the Emphaticness and language and dialect for Voice Onset Time (VOT) in Modern Standard Arabic (MSA)

2013-05-13 · Sulaiman S. AlDahri

The signal sound contains many different features, including Voice Onset Time (VOT), which is a very important feature of stop sounds in many languages. The only application of VOT values is stopping phoneme subsets. Thi…

Universal Phone Recognition with a Multilingual Allophone System

2020-02-26 · Xinjian Li, Siddharth Dalmia, Juncheng Li, Matthew Lee 외

Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonem…

speech-recognitionSpeech Recognition

Automatic Discrimination between Inherited and Borrowed Latin Words in Romance Languages

2021-11-01 · Findings (EMNLP) 2021 11 · Alina Maria Cristea, Liviu P. Dinu, Simona Georgescu, Mihnea-Lucian Mihai 외

In this paper, we address the problem of automatically discriminating between inherited and borrowed Latin words. We introduce a new dataset and investigate the case of Romance languages (Romanian, Italian, French, Spani…