paper-with-me

홈 › Papers

Assessing Dutch Syllabification Algorithms and Improving Accuracy by Combining Phonetic and Orthographic Information through Deep Learning

2026-04-10 · Gus Lathouwers, Wieke Harmsen, Catia Cucchiarini, Helmer Strik arxiv

Syllabification describes the task of dividing words into syllables. Due to many rules and exceptions, training an algorithm to perform syllabification with high accuracy remains a challenge. Throughout the last decades, different algorithms have been put forth for Dutch syllabification, yet a comprehensive comparative assessment has not been done. Additionally, deep learning has gained significant popularity within NLP in recent years, yet no modern deep-learning based framework has been developed for Dutch orthographic syllabification. Finally, phonetic and orthographic syllabification algorithms have been examined separately, but not in combination. The aim of the current research was twofold: (a) to examine the performance of existing Dutch syllabification algorithms, and (b) to investigate whether combining phonetic and orthographic information into a single model can increase syllabification performance. To compare the performance of algorithms, four algorithms (Brandt Corstius, Liang, Trogkanis-Elkan (CRF), and a newly conceived deep-learning model) were applied to three different datasets (dictionary words, loanwords, pseudowords). The algorithms show varying performance across datasets, with the data-driven algorithms outperforming a knowledge-based algorithm in all but one condition. The new deep-learning methods developed led to increased performance compared to the best found in the literature (99.65% word accuracy, a 0.14% improvement). An analysis of the words for which adding phonetic information improved syllabification performance indicates that these were words in which the orthographic ambiguity could be resolved by information on pronunciation. Future research could examine other areas where phonetic information can benefit orthographic processing. In addition, the newly developed deep learning frameworks can be applied to other languages than Dutch.

📄 PDF Abstract BibTeX arXiv:2605.28834

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language-Agnostic Syllabification with Neural Sequence Labeling

2019-09-29 · Jacob Krantz, Maxwell Dulin, Paul De Palma

The identification of syllables within phonetic sequences is known as syllabification. This task is thought to play an important role in natural language understanding, speech production, and the development of speech re…

Chunkingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+7

Design and Implementation of a Tool for Extracting Uzbek Syllables

2023-12-25 · Ulugbek Salaev, Elmurod Kuriyozov, Gayrat Matlatipov

The accurate syllabification of words plays a vital role in various Natural Language Processing applications. Syllabification is a versatile linguistic tool with applications in linguistic research, language technology, …

Automatic Syllabification for Manipuri language

2016-12-01 · COLING 2016 12 · Loitongbam Gyanendro Singh, Lenin Laitonjam, Sanasam Ranbir Singh

Development of hand crafted rule for syllabifying words of a language is an expensive task. This paper proposes several data-driven methods for automatic syllabification of words written in Manipuri language. Manipuri is…

Automatic Speech Recognition (ASR)SegmentationSpeech RecognitionSpeech Synthesis+1

Tenyidie Syllabification corpus creation and deep learning applications

2025-10-01 · Teisovi Angami, Kevisino Khate arxiv

The Tenyidie language is a low-resource language of the Tibeto-Burman family spoken by the Tenyimia Community of Nagaland in the north-eastern part of India and is considered a major language in Nagaland. It is tonal, Su…

Part-Of-Speech TaggingMachine Translation

Syllabification of the Divine Comedy

2020-10-26 · Andrea Asperti, Stefano Dal Bianco

We provide a syllabification algorithm for the Divine Comedy using techniques from probabilistic and constraint programming. We particularly focus on the synalephe, addressed in terms of the "propensity" of a word to tak…

Clustering