Integration of Japanese Papers Into the DBLP Data Set
If someone is looking for a certain publication in the field of computer science, the searching person is likely to use the DBLP to find the desired publication. The DBLP data set is continuously extended with new publications, or rather their metadata, for example the names of involved authors, the title and the publication date. While the size of the data set is already remarkable, specific areas can still be improved. The DBLP offers a huge collection of English papers because most papers concerning computer science are published in English. Nevertheless, there are official publications in other languages which are supposed to be added to the data set. One kind of these are Japanese papers. This diploma thesis will show a way to automatically process publication lists of Japanese papers and to make them ready for an import into the DBLP data set. Especially important are the problems along the way of processing, such as transcription handling and Personal Name Matching with Japanese names.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science Research
DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and e…
ArticlesAn Instance-based Plus Ensemble Learning Method for Classification of Scientific Papers
The exponential growth of scientific publications in recent years has posed a significant challenge in effective and efficient categorization. This paper introduces a novel approach that combines instance-based learning …
ClassificationEnsemble LearningBERTologyNavigator: Advanced Question Answering with BERT-based Semantics
The development and integration of knowledge graphs and language models has significance in artificial intelligence and natural language processing. In this study, we introduce the BERTologyNavigator -- a two-phased syst…
Knowledge GraphsNavigateQuestion AnsweringRelation+1Linked Papers With Code: The Latest in Machine Learning as an RDF Knowledge Graph
In this paper, we introduce Linked Papers With Code (LPWC), an RDF knowledge graph that provides comprehensive, current information about almost 400,000 machine learning publications. This includes the tasks addressed, t…
Knowledge Graph EmbeddingsNLQxform: A Language Model-based Question to SPARQL Transformer
In recent years, scholarly data has grown dramatically in terms of both scale and complexity. It becomes increasingly challenging to retrieve information from scholarly knowledge graphs that include large-scale heterogen…
Graph Question AnsweringKnowledge GraphsLanguage ModelingLanguage Modelling+1