Identifying and Improving Dataset References in Social Sciences Full Texts
Scientific full text papers are usually stored in separate places than their underlying research datasets. Authors typically make references to datasets by mentioning them for example by using their titles and the year of publication. However, in most cases explicit links that would provide readers with direct access to referenced datasets are missing. Manually detecting references to datasets in papers is time consuming and requires an expert in the domain of the paper. In order to make explicit all links to datasets in papers that have been published already, we suggest and evaluate a semi-automatic approach for finding references to datasets in social sciences papers. Our approach does not need a corpus of papers (no cold start problem) and it performs well on a small test corpus (gold standard). Our approach achieved an F-measure of 0.84 for identifying references in full texts and an F-measure of 0.83 for finding correct matches of detected references in the da|ra dataset registry.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Machine learning in the social and health sciences
The uptake of machine learning (ML) approaches in the social and health sciences has been rather slow, and research using ML for social and health research questions remains fragmented. This may be due to the separate de…
BIG-bench Machine LearningCausal InferenceCharacterizing References from Different Disciplines: A Perspective of Citation Content Analysis
Multidisciplinary cooperation is now common in research since social issues inevitably involve multiple disciplines. In research articles, reference information, especially citation content, is an important representatio…
ArticlesHow to foster innovation in the social sciences? Qualitative evidence from focus group workshops at Oxford University
This report addresses challenges and opportunities for innovation in the social sciences at the University of Oxford. It summarises findings from two focus group workshops with innovation experts from the University ecos…
A Semantic Approach for User-Brand Targeting in On-Line Social Networks
We propose a general framework for the recommendation of possible customers (users) to advertisers (e.g., brands) based on the comparison between On-line Social Network profiles. In particular, we represent both user and…
SsciBERT: A Pre-trained Language Model for Social Science Texts
The academic literature of social sciences records human civilization and studies human social problems. With its large-scale growth, the ways to quickly find existing research on relevant issues have become an urgent de…
Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+1