Cleaning English Abstracts of Scientific Publications
Scientific abstracts are often used as proxies for the content and thematic focus of research publications. However, a significant share of published abstracts contains extraneous information-such as publisher copyright statements, section headings, author notes, registrations, and bibliometric or bibliographic metadata-that can distort downstream analyses, particularly those involving document similarity or textual embeddings. We introduce an open-source, easy-to-integrate language model designed to clean English-language scientific abstracts by automatically identifying and removing such clutter. We demonstrate that our model is both conservative and precise, alters similarity rankings of cleaned abstracts and improves information content of standard-length embeddings.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Scielo Corpus: a Parallel Corpus of Scientific Publications for Biomedicine
The biomedical scientific literature is a rich source of information not only in the English language, for which it is more abundant, but also in other languages, such as Portuguese, Spanish and French. We present the fi…
Machine TranslationTranslationJoint Content-Context Analysis of Scientific Publications: Identifying Opportunities for Collaboration in Cognitive Science
This work studies publications in the field of cognitive science and utilizes mathematical techniques to connect the analysis of the papers' content (abstracts) to the context (citation, journals). We apply hierarchical …
Community DetectionJoint Content-Context Analysis of Scientific Publications: Identifying Opportunities for Collaboration in Cognitive Science
This work studies publications in cognitive science and utilizes natural language processing and graph theoretical techniques to connect the analysis of the papers' content (abstracts) to the context (citation, journals)…
Community DetectionHow scientific literature has been evolving over the time? A novel statistical approach using tracking verbal-based methods
This paper provides a global vision of the scientific publications related with the Systemic Lupus Erythematosus (SLE), taking as starting point abstracts of articles. Through the time, abstracts have been evolving towar…
ArticlesTagging Scientific Publications using Wikipedia and Natural Language Processing Tools. Comparison on the ArXiv Dataset
In this work, we compare two simple methods of tagging scientific publications with labels reflecting their content. As a first source of labels Wikipedia is employed, second label set is constructed from the noun phrase…
BIG-bench Machine LearningClustering