paper-with-me

Papers

Corpus analysis without prior linguistic knowledge - unsupervised mining of phrases and subphrase structure

2016-02-18 · Stefan Gerdjikov, Klaus U. Schulz

When looking at the structure of natural language, "phrases" and "words" are central notions. We consider the problem of identifying such "meaningful subparts" of language of any length and underlying composition principles in a completely corpus-based and language-independent way without using any kind of prior linguistic knowledge. Unsupervised methods for identifying "phrases", mining subphrase structure and finding words in a fully automated way are described. This can be considered as a step towards automatically computing a "general dictionary and grammar of the corpus". We hope that in the long run variants of our approach turn out to be useful for other kind of sequence data as well, such as, e.g., speech, genom sequences, or music annotation. Even if we are not primarily interested in immediate applications, results obtained for a variety of languages show that our methods are interesting for many practical tasks in text mining, terminology extraction and lexicography, search engine technology, and related fields.

📄 PDF Abstract BibTeX arXiv:1602.05772

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rich Character-Level Information for Korean Morphological Analysis and Part-of-Speech Tagging

2018-06-28 · COLING 2018 8 · Andrew Matteson, Chanhee Lee, Young-Bum Kim, Heuiseok Lim

Due to the fact that Korean is a highly agglutinative, character-rich language, previous work on Korean morphological analysis typically employs the use of sub-character features known as graphemes or otherwise utilizes …

Morphological AnalysisPart-Of-Speech TaggingSentence

How Universal is Metonymy? Results from a Large-Scale Multilingual Analysis

2022-07-01 · NAACL (SIGTYP) 2022 7 · Temuulen Khishigsuren, Gábor Bella, Thomas Brochhagen, Daariimaa Marav 외

Metonymy is regarded by most linguists as a universal cognitive phenomenon, especially since the emergence of the theory of conceptual mappings. However, the field data backing up claims of universality has not been larg…

A Domain and Language Independent Named Entity Classification Approach Based on Profiles and Local Information

2017-09-01 · RANLP 2017 9 · Isabel Moreno, Mar{\'\i}a Teresa Rom{\'a}-Ferri, Paloma Moreda Pozo

This paper presents a Named Entity Classification system, which employs machine learning. Our methodology employs local entity information and profiles as feature set. All features are generated in an unsupervised manner…

General ClassificationNamed Entity Recognition (NER)Question AnsweringText Generation+1

Modeling the Readability of German Targeting Adults and Children: An empirically broad analysis and its cross-corpus validation

2018-08-01 · COLING 2018 8 · Zarah Wei{\ss}, Detmar Meurers

We analyze two novel data sets of German educational media texts targeting adults and children. The analysis is based on 400 automatically extracted measures of linguistic complexity from a wide range of linguistic domai…

Binary ClassificationCross-corpusInformation RetrievalText Simplification

Visual Relationship Detection with Language prior and Softmax

2019-04-16 · Jaewon Jung, Jongyoul Park

Visual relationship detection is an intermediate image understanding task that detects two objects and classifies a predicate that explains the relationship between two objects in an image. The three components are lingu…

Knowledge DistillationRelationship DetectionVisual Relationship Detection