paper-with-me

Papers

Constrained Sequence-to-sequence Semitic Root Extraction for Enriching Word Embeddings

2019-08-01 · WS 2019 8 · Ahmed El-Kishky, Xingyu Fu, Aseel Addawood, Nahil Sobh, Clare Voss, Jiawei Han

In this paper, we tackle the problem of {``}root extraction{''} from words in the Semitic language family. A challenge in applying natural language processing techniques to these languages is the data sparsity problem that arises from their rich internal morphology, where the substructure is inherently non-concatenative and morphemes are interdigitated in word formation. While previous automated methods have relied on human-curated rules or multiclass classification, they have not fully leveraged the various combinations of regular, sequential concatenative morphology within the words and the internal interleaving within templatic stems of roots and patterns. To address this, we propose a constrained sequence-to-sequence root extraction method. Experimental results show our constrained model outperforms a variety of methods at root extraction. Furthermore, by enriching word embeddings with resulting decompositions, we show improved results on word analogy, word similarity, and language modeling tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingWord EmbeddingsWord Similarity

Similar Papers 제목 키워드 기반

Fixing the Infix: Unsupervised Discovery of Root-and-Pattern Morphology

2017-02-07 · Tarek Sakakini, Suma Bhat, Pramod Viswanath

We present an unsupervised and language-agnostic method for learning root-and-pattern morphology in Semitic languages. This form of morphology, abundant in Semitic languages, has not been handled in prior unsupervised ap…

Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

2026-03-26 · A. Feder Cooper, Mark A. Lemley, Christopher De Sa, Lea Duesterwald 외 arxiv

Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences. Probabilistic extraction -- computing the probability of generatin…

Pattern-and-root inflectional morphology: the Arabic broken plural

2026-05-21 · Alexis Amid Neme, Eric Laporte arxiv

We present a substantially implemented model of description of the inflectional morphology of Arabic nouns, with special attention to the management of dictionaries and other language resources by Arabic-speaking linguis…

Extracting N-ary Cross-sentence Relations using Constrained Subsequence Kernel

2020-06-15 · Sachin Pawar, Pushpak Bhattacharyya, Girish K. Palshikar

Most of the past work in relation extraction deals with relations occurring within a sentence and having only two entity arguments. We propose a new formulation of the relation extraction task where the relations are mor…

RelationRelation ExtractionSentencevalid

Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data

2025-09-16 · Kurt Micallef, Nizar Habash, Claudia Borg arxiv

Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin scri…

Machine TranslationData Augmentation