paper-with-me

Papers

Estimation of the Frequency of Occurrence of Italian Phonemes in Text

2021-01-14 · Javi Arango, Alec DeCaprio, Sunwoo Baik, Luca De Nardis, Stefanie Shattuck-Hufnagel, Maria Gabriella Di Benedetto

The purpose of this project was to derive a reliable estimate of the frequency of occurrence of the 30 phonemes - plus consonant geminated counterparts - of the Italian language, based on four selected written texts. Since no comparable dataset was found in previous literature, the present analysis may serve as a reference in future studies. Four textual sources were considered: Come si fa una tesi di laurea: le materie umanistiche by Umberto Eco, I promessi sposi by Alessandro Manzoni, a recent article in Corriere della Sera (a popular daily Italian newspaper), and In altre parole by Jhumpa Lahiri. The sources were chosen to represent varied genres, subject matter, time periods, and writing styles. Results of the analysis, which also included an analysis of variance, showed that, for all four sources, the frequencies of occurrence reached relatively stable values after about 6,000 phonemes (approx. 1,250 words), varying by <0.025%. Estimated frequencies are provided for each single source and as an average across sources.

📄 PDF Abstract BibTeX arXiv:2101.06147

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

FA 설명 없음

Similar Papers 제목 키워드 기반

ITALIC: An Italian Intent Classification Dataset

2023-06-14 · Alkis Koudounas, Moreno La Quatra, Lorenzo Vaiani, Luca Colomba 외

Recent large-scale Spoken Language Understanding datasets focus predominantly on English and do not account for language-specific phenomena such as particular phonemes or words in different lects. We introduce ITALIC, th…

Classificationintent-classificationIntent Classificationspeech-recognition+2

Stochastic model for phonemes uncovers an author-dependency of their usage

2015-10-05 · Weibing Deng, Armen E. Allahverdyan

We study rank-frequency relations for phonemes, the minimal units that still relate to linguistic meaning. We show that these relations can be described by the Dirichlet distribution, a direct analogue of the ideal-gas m…

WAGS: A Beautiful English-Italian Benchmark Supporting Word Alignment Evaluation on Rare Words

2016-05-01 · LREC 2016 5 · Luisa Bentivogli, Mauro Cettolo, M. Amin Farajian, Marcello Federico

This paper presents WAGS (Word Alignment Gold Standard), a novel benchmark which allows extensive evaluation of WA tools on out-of-vocabulary (OOV) and rare words. WAGS is a subset of the Common Test section of the Europ…

SentenceWord Alignment

A Pedagogical Application of NooJ in Language Teaching: The Adjective in Spanish and Italian

2018-08-01 · COLING 2018 8 · Andrea Rodrigo, Mario Monteleone, Silvia Reyes

In this paper, a pedagogical application of NooJ to the teaching and learning of Spanish as a foreign language is presented, which is directed to a specific addressee: learners whose mother tongue is Italian. The categor…

NLP and Public Engagement: The Case of the Italian School Reform

2016-05-01 · LREC 2016 5 · Tommaso Caselli, Giovanni Moretti, Rachele Sprugnoli, Sara Tonelli 외

In this paper we present PIERINO (PIattaforma per l{'}Estrazione e il Recupero di INformazione Online), a system that was implemented in collaboration with the Italian Ministry of Education, University and Research to an…