paper-with-me

Papers

Integrating curation into scientific publishing to train AI models

2023-10-31 · Jorge Abreu-Vicente, Hannah Sonntag, Thomas Eidens, Cassie S. Mitchell, Thomas Lemberger

High throughput extraction and structured labeling of data from academic articles is critical to enable downstream machine learning applications and secondary analyses. We have embedded multimodal data curation into the academic publishing process to annotate segmented figure panels and captions. Natural language processing (NLP) was combined with human-in-the-loop feedback from the original authors to increase annotation accuracy. Annotation included eight classes of bioentities (small molecules, gene products, subcellular components, cell lines, cell types, tissues, organisms, and diseases) plus additional classes delineating the entities' roles in experiment designs and methodologies. The resultant dataset, SourceData-NLP, contains more than 620,000 annotated biomedical entities, curated from 18,689 figures in 3,223 articles in molecular and cell biology. We evaluate the utility of the dataset to train AI models using named-entity recognition, segmentation of figure captions into their constituent panels, and a novel context-dependent semantic task assessing whether an entity is a controlled intervention target or a measurement object. We also illustrate the use of our dataset in performing a multi-modal task for segmenting figures into panel images and their corresponding captions.

📄 PDF Abstract BibTeX arXiv:2310.20440

Code (1)

source-data/soda-data 공식 구현

Tasks

ArticlesEntity LinkingNamed Entity Recognition (NER)

Similar Papers 제목 키워드 기반

Towards a more sustainable academic publishing system

2021-01-18 · Mohsen Kayal, Jane Ballard, Ehsan Kayal

Communicating new scientific discoveries is key to human progress. Yet, this endeavor is hindered by monetary restrictions for publishing one's findings and accessing other scientists' reports. This process is further ex…

CurateGPT: A flexible language-model assisted biocuration tool

2024-10-29 · Harry Caufield, Carlo Kroll, Shawn T O'Neil, Justin T Reese 외

Effective data-driven biomedical discovery requires data curation: a time-consuming process of finding, organizing, distilling, integrating, interpreting, annotating, and validating diverse information into a structured …

Language ModelingLanguage ModellingmodelPhilosophy

HIKMA: Human-Inspired Knowledge by Machine Agents through a Multi-Agent Framework for Semi-Autonomous Scientific Conferences

2025-10-24 · Zain Ul Abideen Tariq, Mahmood Al-Zubaidi, Uzair Shah, Marco Agus 외 arxiv

HIKMA Semi-Autonomous Conference is the first experiment in reimagining scholarly communication through an end-to-end integration of artificial intelligence into the academic publishing and presentation pipeline. This pa…

Could AI change the scientific publishing market once and for all?

2024-01-26 · Wadim Strielkowski

Artificial-intelligence tools in research like ChatGPT are playing an increasingly transformative role in revolutionizing scientific publishing and re-shaping its economic background. They can help academics to tackle su…

All

ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review

2025-10-09 · Gaurav Sahu, Hugo Larochelle, Laurent Charlin, Christopher Pal arxiv

Peer review is the cornerstone of scientific publishing, yet it suffers from inconsistencies, reviewer subjectivity, and scalability challenges. We introduce ReviewerToo, a modular framework for studying and deploying AI…