paper-with-me

홈 › Papers

Using Titles vs. Full-text as Source for Automated Semantic Document Annotation

2017-05-15 · Lukas Galke, Florian Mai, Alan Schelten, Dennis Brunsch, Ansgar Scherp

A significant part of the largest Knowledge Graph today, the Linked Open Data cloud, consists of metadata about documents such as publications, news reports, and other media articles. While the widespread access to the document metadata is a tremendous advancement, it is yet not so easy to assign semantic annotations and organize the documents along semantic concepts. Providing semantic annotations like concepts in SKOS thesauri is a classical research topic, but typically it is conducted on the full-text of the documents. For the first time, we offer a systematic comparison of classification approaches to investigate how far semantic annotations can be conducted using just the metadata of the documents such as titles published as labels on the Linked Open Data cloud. We compare the classifications obtained from analyzing the documents' titles with semantic annotations obtained from analyzing the full-text. Apart from the prominent text classification baselines kNN and SVM, we also compare recent techniques of Learning to Rank and neural networks and revisit the traditional methods logistic regression, Rocchio, and Naive Bayes. The results show that across three of our four datasets, the performance of the classifications using only titles reaches over 90% of the quality compared to the classification performance when using the full-text. Thus, conducting document classification by just using the titles is a reasonable approach for automated semantic annotation and opens up new possibilities for enriching Knowledge Graphs.

📄 PDF Abstract BibTeX arXiv:1705.05311

Code (1)

Quadflor/quadflor 공식 구현 tf

Tasks

ArticlesClassificationDocument ClassificationGeneral ClassificationKnowledge GraphsLearning-To-Ranktext-classificationText Classification

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Building a Hebrew Semantic Role Labeling Lexical Resource from Parallel Movie Subtitles

2020-05-17 · LREC 2020 5 · Ben Eyal, Michael Elhadad

We present a semantic role labeling resource for Hebrew built semi-automatically through annotation projection from English. This corpus is derived from the multilingual OpenSubtitles dataset and includes short informal …

Morphological AnalysisSemantic Role Labeling

TocBERT: Medical Document Structure Extraction Using Bidirectional Transformers

2024-06-27 · Majd Saleh, Sarra Baghdadi, Stéphane Paquelet

Text segmentation holds paramount importance in the field of Natural Language Processing (NLP). It plays an important role in several NLP downstream tasks like information retrieval and document summarization. In this wo…

Document SummarizationHierarchical Text SegmentationInformation Retrievalnamed-entity-recognition+5

Leveraging the Inherent Hierarchy of Vacancy Titles for Automated Job Ontology Expansion

2020-04-06 · LREC 2020 5 · Jeroen Van Hautte, Vincent Schelstraete, Mikaël Wornoo

Machine learning plays an ever-bigger part in online recruitment, powering intelligent matchmaking and job recommendations across many of the world's largest job platforms. However, the main text is rarely enough to full…

NER

Generating Multilingual Parallel Corpus Using Subtitles

2018-04-11 · Farshad Jafari

Neural Machine Translation with its significant results, still has a great problem: lack or absence of parallel corpus for many languages. This article suggests a method for generating considerable amount of parallel cor…

Machine TranslationSentence

Bi-Text Alignment of Movie Subtitles for Spoken English-Arabic Statistical Machine Translation

2016-09-05 · Fahad Al-Obaidli, Stephen Cox, Preslav Nakov

We describe efforts towards getting better resources for English-Arabic machine translation of spoken text. In particular, we look at movie subtitles as a unique, rich resource, as subtitles in one language often get tra…

Machine TranslationTranslation