PubMedAbstractsSubsetEmbedded
홈페이지 · 논문 1편
This dataset contains a probabilistic sample of ~2.4 million PubMed abstracts, enriched with precomputed dense embeddings (title + abstract), from the ncbi/MedCPT-Article-Encoder model. It is derived from public metadata made available via the National Library of Medicine (NLM) and was used in the paper *Efficient and Reproducible Biomedical QA using Retrieval-Augmented Generation*.
Each entry includes:
- title: Title of the publication
- abstract: Abstract content
- PMID: PubMed identifier
- embedding: 768-dimensional float32 vector from MedCPT