paper-with-me

홈 › Papers

Uzbek-English and Turkish-English Morpheme Alignment Corpora

2016-05-01 · LREC 2016 5 · Xuansong Li, Jennifer Tracey, Stephen Grimes, Stephanie Strassel

Morphologically-rich languages pose problems for machine translation (MT) systems, including word-alignment errors, data sparsity and multiple affixes. Current alignment models at word-level do not distinguish words and morphemes, thus yielding low-quality alignment and subsequently affecting end translation quality. Models using morpheme-level alignment can reduce the vocabulary size of morphologically-rich languages and overcomes data sparsity. The alignment data based on smallest units reveals subtle language features and enhances translation quality. Recent research proves such morpheme-level alignment (MA) data to be valuable linguistic resources for SMT, particularly for languages with rich morphology. In support of this research trend, the Linguistic Data Consortium (LDC) created Uzbek-English and Turkish-English alignment data which are manually aligned at the morpheme level. This paper describes the creation of MA corpora, including alignment and tagging process and approaches, highlighting annotation challenges and specific features of languages with rich morphology. The light tagging annotation on the alignment layer adds extra value to the MA data, facilitating users in flexibly tailoring the data for various MT model training.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslationWord Alignment

Similar Papers 제목 키워드 기반

VerChol -- Grammar-First Tokenization for Agglutinative Languages

2026-03-06 · Prabhu Raja arxiv

Tokenization is the foundational step in all large language model (LLM) pipelines, yet the dominant approach Byte Pair Encoding (BPE) and its variants is inherently script agnostic and optimized for English like morpholo…

AMR Alignment for Morphologically-rich and Pro-drop Languages

2022-05-01 · ACL 2022 5 · K. Elif Oral, Gülşen Eryiğit

Alignment between concepts in an abstract meaning representation (AMR) graph and the words within a sentence is one of the important stages of AMR parsing. Although there exist high performing AMR aligners for English, u…

Abstract Meaning RepresentationAMR ParsingSentence

Developing a Comprehensive Framework for Sentiment Analysis in Turkish

2025-11-29 · Cem Rifki Aydin arxiv

In this thesis, we developed a comprehensive framework for sentiment analysis that takes its many aspects into account mainly for Turkish. We have also proposed several approaches specific to sentiment analysis in Englis…

Sentiment AnalysisTerm Extraction

Cross-Lingual Word Embeddings for Turkic Languages

2020-05-17 · LREC 2020 5 · Elmurod Kuriyozov, Yerai Doval, Carlos Gómez-Rodríguez

There has been an increasing interest in learning cross-lingual word embeddings to transfer knowledge obtained from a resource-rich language, such as English, to lower-resource languages for which annotated data is scarc…

Cross-Lingual Word EmbeddingsSentiment AnalysisWord Embeddings

Using Interlinear Glosses as Pivot in Low-Resource Multilingual Machine Translation

2019-11-07 · Zhong Zhou, Lori Levin, David R. Mortensen, Alex Waibel

We demonstrate a new approach to Neural Machine Translation (NMT) for low-resource languages using a ubiquitous linguistic resource, Interlinear Glossed Text (IGT). IGT represents a non-English sentence as a sequence of …

Machine TranslationNMTSentenceTranslation