paper-with-me

홈 › Papers

[Citation needed] Data usage and citation practices in medical imaging conferences

2024-02-05 · Théo Sourget, Ahmet Akkoç, Stinna Winther, Christine Lyngbye Galsgaard, Amelia Jiménez-Sánchez, Dovile Juodelyte, Caroline Petitjean, Veronika Cheplygina

Medical imaging papers often focus on methodology, but the quality of the algorithms and the validity of the conclusions are highly dependent on the datasets used. As creating datasets requires a lot of effort, researchers often use publicly available datasets, there is however no adopted standard for citing the datasets used in scientific papers, leading to difficulty in tracking dataset usage. In this work, we present two open-source tools we created that could help with the detection of dataset usage, a pipeline \url{https://github.com/TheoSourget/Public_Medical_Datasets_References} using OpenAlex and full-text analysis, and a PDF annotation software \url{https://github.com/TheoSourget/pdf_annotator} used in our study to manually label the presence of datasets. We applied both tools on a study of the usage of 20 publicly available medical datasets in papers from MICCAI and MIDL. We compute the proportion and the evolution between 2013 and 2023 of 3 types of presence in a paper: cited, mentioned in the full text, cited and mentioned. Our findings demonstrate the concentration of the usage of a limited set of datasets. We also highlight different citing practices, making the automation of tracking difficult.

📄 PDF Abstract BibTeX arXiv:2402.03003

Code (2)

theosourget/pdf_annotator 공식 구현
theosourget/public_medical_datasets_references 공식 구현

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

Assigning credit to scientific datasets using article citation networks

2020-01-16 · Tong Zeng, Longfeng Wu, Sarah Bratt, Daniel E. Acuna

A citation is a well-established mechanism for connecting scientific artifacts. Citation networks are used by citation analysis for a variety of reasons, prominently to give credit to scientists' work. However, because o…

How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?

2025-04-03 · Andres Algaba, Vincent Holst, Floriano Tori, Melika Mobini 외

The spread of scientific knowledge depends on how researchers discover and cite previous work. The adoption of large language models (LLMs) in the scientific research process introduces a new layer to these citation prac…

scientific discovery

Cross-Lingual Citations in English Papers: A Large-Scale Analysis of Prevalence, Usage, and Impact

2021-11-07 · Tarek Saier, Michael Färber, Tornike Tsereteli

Citation information in scholarly data is an important source of insight into the reception of publications and the scholarly discourse. Outcomes of citation analyses and the applicability of citation based machine learn…

Citation Intent ClassificationCross-Lingual Entity LinkingData Visualization

Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias

2024-05-24 · Andres Algaba, Carmen Mazijn, Vincent Holst, Floriano Tori 외

Citation practices are crucial in shaping the structure of scientific knowledge, yet they are often influenced by contemporary norms and biases. The emergence of Large Language Models (LLMs) introduces a new dynamic to t…

Retrieval-augmented Generation

Quantifying and suppressing ranking bias in a large citation network

2017-03-23 · Vaccario Giacomo, Medo Matus, Wider Nicolas, Mariani Manuel Sebastian

It is widely recognized that citation counts for papers from different fields cannot be directly compared because different scientific fields adopt different citation practices. Citation counts are also strongly biased b…