paper-with-me

홈 › Papers

Compositional Morpheme Embeddings with Affixes as Functions and Stems as Arguments

2018-07-01 · WS 2018 7 · Daniel Edmiston, Karl Stratos

This work introduces a novel, linguistically motivated architecture for composing morphemes to derive word embeddings. The principal novelty in the work is to treat stems as vectors and affixes as functions over vectors. In this way, our model{'}s architecture more closely resembles the compositionality of morphemes in natural language. Such a model stands in opposition to models which treat morphemes uniformly, making no distinction between stem and affix. We run this new architecture on a dependency parsing task in Korean{---}a language rich in derivational morphology{---}and compare it against a lexical baseline,along with other sub-word architectures. StAffNet, the name of our architecture, shows competitive performance with the state-of-the-art on this task.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dependency ParsingWord Embeddings

Similar Papers 제목 키워드 기반

Bilingual Lexicon Extraction at the Morpheme Level Using Distributional Analysis

2016-05-01 · LREC 2016 5 · Amir Hazem, B{\'e}atrice Daille

Bilingual lexicon extraction from comparable corpora is usually based on distributional methods when dealing with single word terms (SWT). These methods often treat SWT as single tokens without considering their composit…

Translation

PACUTE: Phonology-, Affix-, and Character-level Understanding of Tokens for Filipino

2026-06-13 · Jann Railey Montalan, David Demitri Africa, Jimson Paulo Layacan, Richell Isaiah Flores 외 arxiv

Large language models (LLMs) process text as sequences of subword tokens, which can obscure the character-level and morphological structure that underlies word formation. This limitation is most acute for languages with …

Uzbek-English and Turkish-English Morpheme Alignment Corpora

2016-05-01 · LREC 2016 5 · Xuansong Li, Jennifer Tracey, Stephen Grimes, Stephanie Strassel

Morphologically-rich languages pose problems for machine translation (MT) systems, including word-alignment errors, data sparsity and multiple affixes. Current alignment models at word-level do not distinguish words and …

Machine TranslationTranslationWord Alignment

CWoMP: Morpheme Representation Learning for Interlinear Glossing

2026-03-18 · Morris Alper, Enora Rice, Bhargav Shandilya, Alexis Palmer 외 arxiv

Interlinear glossed text (IGT) is a standard notation for language documentation which is linguistically rich but laborious to produce manually. Recent automated IGT methods treat glosses as character sequences, neglecti…

Representation Learning

Linguistically inspired morphological inflection with a sequence to sequence model

2020-09-04 · Eleni Metheniti, Guenter Neumann, Josef van Genabith

Inflection is an essential part of every human language's morphology, yet little effort has been made to unify linguistic theory and computational methods in recent years. Methods of string manipulation are used to infer…

Language AcquisitionLEMMAMorphological Inflection