Original-Transcribed Text Alignment for Manyosyu Written by Old Japanese Language
We are constructing an annotated diachronic corpora of the Japanese language. In part of thiswork, we construct a corpus of Manyosyu, which is an old Japanese poetry anthology. In thispaper, we describe how to align the transcribed text and its original text semiautomatically to beable to cross-reference them in our Manyosyu corpus. Although we align the original charactersto the transcribed words manually, we preliminarily align the transcribed and original charactersby using an unsupervised automatic alignment technique of statistical machine translation toalleviate the work. We found that automatic alignment achieves an F1-measure of 0.83; thus, each poem has 1{--}2 alignment errors. However, finding these errors and modifying them are less workintensiveand more efficient than fully manual annotation. The alignment probabilities can beutilized in this modification. Moreover, we found that we can locate the uncertain transcriptionsin our corpus and compare them to other transcriptions, by using the alignment probabilities.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
The COPLE2 corpus: a learner corpus for Portuguese
We present the COPLE2 corpus, a learner corpus of Portuguese that includes written and spoken texts produced by learners of Portuguese as a second or foreign language. The corpus includes at the moment a total of 182,474…
LemmatizationPOSUnsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition
An unsupervised text-to-speech synthesis (TTS) system learns to generate speech waveforms corresponding to any written sentence in a language by observing: 1) a collection of untranscribed speech waveforms in that langua…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Sentencespeech-recognition+5Semi-automatically Alignment of Predicates between Speech and OntoNotes data
Speech data currently receives a growing attention and is an important source of information. We still lack suitable corpora of transcribed speech annotated with semantic roles that can be used for semantic role labeling…
Semantic Role LabelingSentenceOpen Set Classification of Untranscribed Handwritten Documents
Huge amounts of digital page images of important manuscripts are preserved in archives worldwide. The amounts are so large that it is generally unfeasible for archivists to adequately tag most of the documents with the r…
Classificationopen-set classificationTAGDevelopment of Natural Language Processing Tools for Cook Islands M\=aori
This paper presents three ongoing projects for NLP in Cook Islands Maori: Untrained Forced Alignment (approx. 9{\%} error when detecting the center of words), speech-to-text (37{\%} WER in the best trained models) and PO…
Machine TranslationPart-Of-Speech TaggingPOSPOS Tagging+2