paper-with-me

홈 › Papers

Utilizing constituent structure for compound analysis

2014-05-01 · LREC 2014 5 · Krist{\'\i}n Bjarnad{\'o}ttir, J{\'o}n Da{\dh}ason

Compounding is extremely productive in Icelandic and multi-word compounds are common. The likelihood of finding previously unseen compounds in texts is thus very high, which makes out-of-vocabulary words a problem in the use of NLP tools. The tool de-scribed in this paper splits Icelandic compounds and shows their binary constituent structure. The probability of a constituent in an unknown (or unanalysed) compound forming a combined constituent with either of its neighbours is estimated, with the use of data on the constituent structure of over 240 thousand compounds from the Database of Modern Icelandic Inflection, and word frequencies from {\'I}slenskur or{\dh}asj{\'o}{\dh}ur, a corpus of approx. 550 million words. Thus, the structure of an unknown compound is derived by com-parison with compounds with partially the same constituents and similar structure in the training data. The granularity of the split re-turned by the decompounder is important in tasks such as semantic analysis or machine translation, where a flat (non-structured) se-quence of constituents is insufficient.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMachine TranslationPart-Of-Speech TaggingSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Kvistur 2.0: a BiLSTM Compound Splitter for Icelandic

2020-04-16 · LREC 2020 5 · Jón Friðrik Daðason, David Erik Mollberg, Hrafn Loftsson, Kristín Bjarnadóttir

In this paper, we present a character-based BiLSTM model for splitting Icelandic compound words, and show how varying amounts of training data affects the performance of the model. Compounding is highly productive in Ice…

Part-Of-Speech Tagging

GhoSt-NN: A Representative Gold Standard of German Noun-Noun Compounds

2016-05-01 · LREC 2016 5 · Sabine Schulte im Walde, Anna H{\"a}tty, Stefan Bott, Nana Khvtisavrishvili

This paper presents a novel gold standard of German noun-noun compounds (Ghost-NN) including 868 compounds annotated with corpus frequencies of the compounds and their constituents, productivity and ambiguity of the cons…

A Psycholinguistic Analysis of BERT's Representations of Compounds

2023-02-14 · Lars Buijtelaar, Sandro Pezzelle

This work studies the semantic representations learned by BERT for compounds, that is, expressions such as sunlight or bodyguard. We build on recent studies that explore semantic information in Transformers at the word l…

Variants of Vector Space Reductions for Predicting the Compositionality of English Noun Compounds

2020-05-01 · LREC 2020 5 · Pegah Alipoor, Sabine Schulte im Walde

Predicting the degree of compositionality of noun compounds such as {``}snowball{''} and {``}butterfly{''} is a crucial ingredient for lexicography and Natural Language Processing applications, to know whether the compou…

Sense disambiguation of compound constituents

2022-04-01 · Carlo Schackow, Stefan Conrad, Ingo Plag

In distributional semantic accounts of the meaning of noun-noun compounds (e.g. starfish, bank account, houseboat) the important role of constituent polysemy remains largely unaddressed(cf. the meaning of star in starfis…

Word Sense Disambiguation