paper-with-me

Papers

Refining Word Segmentation Using a Manually Aligned Corpus for Statistical Machine Translation

2014-10-01 · EMNLP 2014 10 · Xiaolin Wang, Masao Utiyama, Andrew Finch, Eiichiro Sumita
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslationWord Alignment

Similar Papers 제목 키워드 기반

Substring Frequency Features for Segmentation of Japanese Katakana Words with Unlabeled Corpora

2017-11-01 · IJCNLP 2017 11 · Yoshinari Fujinuma, Alvin Grissom II

Word segmentation is crucial in natural language processing tasks for unsegmented languages. In Japanese, many out-of-vocabulary words appear in the phonetic syllabary katakana, making segmentation more difficult due to …

Information RetrievalMachine TranslationSegmentation

LeConTra: A Learner Corpus of English-to-Dutch News Translation

2022-06-01 · LREC 2022 6 · Bram Vanroy, Lieve Macken

We present LeConTra, a learner corpus consisting of English-to-Dutch news translations enriched with translation process data. Three students of a Master’s programme in Translation were asked to translate 50 different En…

SentenceSentence segmentationTranslation

PEACH: A sentence-aligned Parallel English-Arabic Corpus for Healthcare

2025-08-07 · Rania Al-Sabbagh arxiv

This paper introduces PEACH, a sentence-aligned parallel English-Arabic corpus of healthcare texts encompassing patient information leaflets and educational materials. The corpus contains 51,671 parallel sentences, total…

Machine Translation

Khmer Word Segmentation Using Conditional Random Fields

2015-10-15 · Vichet Chea, Ye Kyaw Thu, Chenchen Ding, Masao Utiyama 외

Word Segmentation is a critical task that is the foundation of much natural language processing research. This paper is a study of Khmer word segmentation using an approach based on conditional random fields (CRFs). A…

SegmentationText SegmentationTranslation

The Parallel Meaning Bank: Towards a Multilingual Corpus of Translations Annotated with Compositional Meaning Representations

2017-02-13 · EACL 2017 4 · Lasha Abzianidze, Johannes Bjerva, Kilian Evang, Hessel Haagsma 외

The Parallel Meaning Bank is a corpus of translations annotated with shared, formal meaning representations comprising over 11 million words divided over four languages (English, German, Italian, and Dutch). Our approach…