Text2Cohort: Facilitating Intuitive Access to Biomedical Data with Natural Language Cohort Discovery
The Imaging Data Commons (IDC) is a cloud-based database that provides researchers with open access to cancer imaging data, with the goal of facilitating collaboration. However, cohort discovery within the IDC database has a significant technical learning curve. Recently, large language models (LLM) have demonstrated exceptional utility for natural language processing tasks. We developed Text2Cohort, a LLM-powered toolkit to facilitate user-friendly natural language cohort discovery in the IDC. Our method translates user input into IDC queries using grounding techniques and returns the query's response. We evaluate Text2Cohort on 50 natural language inputs, from information extraction to cohort discovery. Our toolkit successfully generated responses with an 88% accuracy and 0.94 F1 score. We demonstrate that Text2Cohort can enable researchers to discover and curate cohorts on IDC with high levels of accuracy using natural language in a more intuitive and user-friendly way.
Code (1)
Tasks
Language ModellingLarge Language ModelPrompt EngineeringSimilar Papers 제목 키워드 기반
Integrative Factorization of Bidimensionally Linked Matrices
Advances in molecular "omics'" technologies have motivated new methodology for the integration of multiple sources of high-content biomedical data. However, most statistical methods for integrating multiple data matrices…
Dimensionality ReductionMissing ValuesA Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI
Despite the excitement behind biomedical artificial intelligence (AI), access to high-quality, diverse, and large-scale data - the foundation for modern AI systems - is still a bottleneck to unlocking its full potential.…
ArticlesCohortFinder: an open-source tool for data-driven partitioning of biomedical image cohorts to yield robust machine learning models
Batch effects (BEs) refer to systematic technical differences in data collection unrelated to biological variations whose noise is shown to negatively impact machine learning (ML) model generalizability. Here we release …
Adaptive Recruitment Resource Allocation to Improve Cohort Representativeness in Participatory Biomedical Datasets
Large participatory biomedical studies, studies that recruit individuals to join a dataset, are gaining popularity and investment, especially for analysis by modern AI methods. Because they purposively recruit participan…
Bizard: A Community-Driven Platform for Accelerating and Enhancing Biomedical Data Visualization
Bizard is a novel visualization code repository designed to simplify data analysis in biomedical research. It integrates diverse visualization codes, facilitating the selection and customization of optimal visualization …
Data Visualization