Refining Word Segmentation Using a Manually Aligned Corpus for Statistical Machine Translation
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationWord AlignmentSimilar Papers 제목 키워드 기반
Substring Frequency Features for Segmentation of Japanese Katakana Words with Unlabeled Corpora
Word segmentation is crucial in natural language processing tasks for unsegmented languages. In Japanese, many out-of-vocabulary words appear in the phonetic syllabary katakana, making segmentation more difficult due to …
Information RetrievalMachine TranslationSegmentationLeConTra: A Learner Corpus of English-to-Dutch News Translation
We present LeConTra, a learner corpus consisting of English-to-Dutch news translations enriched with translation process data. Three students of a Master’s programme in Translation were asked to translate 50 different En…
SentenceSentence segmentationTranslationPEACH: A sentence-aligned Parallel English-Arabic Corpus for Healthcare
This paper introduces PEACH, a sentence-aligned parallel English-Arabic corpus of healthcare texts encompassing patient information leaflets and educational materials. The corpus contains 51,671 parallel sentences, total…
Machine TranslationKhmer Word Segmentation Using Conditional Random Fields
Word Segmentation is a critical task that is the foundation of much natural language processing research. This paper is a study of Khmer word segmentation using an approach based on conditional random fields (CRFs). A…
SegmentationText SegmentationTranslationThe Parallel Meaning Bank: Towards a Multilingual Corpus of Translations Annotated with Compositional Meaning Representations
The Parallel Meaning Bank is a corpus of translations annotated with shared, formal meaning representations comprising over 11 million words divided over four languages (English, German, Italian, and Dutch). Our approach…