paper-with-me

홈 › Papers

Robust Biomedical Publication Type and Study Design Classification with Knowledge-Guided Perturbations

2026-05-12 · Shufan Ming, Joe D. Menke, Neil R. Smalheiser, Halil Kilicoglu arxiv

Accurately and consistently indexing biomedical literature by publication type and study design is essential for supporting evidence synthesis and knowledge discovery. Prior work on automated publication type and study design indexing has primarily focused on expanding label coverage, enriching feature representations, and improving in-domain accuracy, with evaluation typically conducted on data drawn from the same distribution as training. Although pretrained biomedical language models achieve strong performance under these settings, models optimized for in-domain accuracy may rely on superficial lexical or dataset-specific cues, resulting in reduced robustness under distributional shift. In this study, we introduce an evaluation framework based on controlled semantic perturbations to assess the robustness of a publication type classifier and investigate robustness-oriented training strategies that combine entity masking and domain-adversarial training to mitigate reliance on spurious topical correlations. Our results show that the commonly observed trade-off between robustness and in-domain accuracy can be mitigated when robustness objectives are designed to selectively suppress non-task-defining features while preserving salient methodological signals. We find that these improvements arise from two complementary mechanisms: (1) increased reliance on explicit methodological cues when such cues are present in the input, and (2) reduced reliance on spurious domain-specific topical features. These findings highlight the importance of feature-level robustness analysis for publication type and study design classification and suggest that refining masking and adversarial objectives to more selectively suppress topical information may further improve robustness. Data, code, and models are available at: https://github.com/ScienceNLP-Lab/MultiTagger-v2/tree/main/ICHI

📄 PDF Abstract BibTeX arXiv:2605.11502

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Divergent Characteristics of Biomedical Research across Publication Types: A Quantitative Analysis on the Aging-related Research

2024-01-09 · Chenxing Qian, Qingyue Guo

This paper investigates differences in characteristics across publication types for aging-related genetic research. We utilized bibliometric data for five model species retrieved from authoritative databases including Pu…

Y-Mol: A Multiscale Biomedical Knowledge-Guided Large Language Model for Drug Development

2024-10-15 · Tengfei Ma, Xuan Lin, Tianle Li, Chaoyi Li 외

Large Language Models (LLMs) have recently demonstrated remarkable performance in general tasks across various fields. However, their effectiveness within specific domains such as drug development remains challenges. To …

Drug DesignKnowledge GraphsLanguage ModelingLanguage Modelling+1

A PubMed-Scale Dataset of Structured Biomedical Abstracts

2026-06-09 · Chia-Hsuan Chang, Haerin Song, Brian Ondov, Hua Xu arxiv

Structured abstracts are important for biomedical literature processing, by facilitating information retrieval, text mining, and knowledge synthesis. However, a vast portion of abstracts indexed in PubMed remain unstruct…

Information ExtractionInformation Retrieval

Bf3R at SemEval-2018 Task 7: Evaluating Two Relation Extraction Tools for Finding Semantic Relations in Biomedical Abstracts

2018-06-01 · SEMEVAL 2018 6 · Mariana Neves, Daniel Butzke, Gilbert Sch{\"o}nfelder, Barbara Grune

Automatic extraction of semantic relations from text can support finding relevant information from scientific publications. We describe our participation in Task 7 of SemEval-2018 for which we experimented with two relat…

Information RetrievalRelation Extraction

In-domain Context-aware Token Embeddings Improve Biomedical Named Entity Recognition

2018-10-01 · WS 2018 10 · Golnar Sheikhshabbafghi, Inanc Birol, Anoop Sarkar

Rapidly expanding volume of publications in the biomedical domain makes it increasingly difficult for a timely evaluation of the latest literature. That, along with a push for automated evaluation of clinical reports, pr…

Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+5