paper-with-me

Papers

Model-assisted cohort selection with bias analysis for generating large-scale cohorts from the EHR for oncology research

2020-01-13 · Benjamin Birnbaum, Nathan Nussbaum, Katharina Seidl-Rathkopf, Monica Agrawal, Melissa Estevez, Evan Estola, Joshua Haimson, Lucy He, Peter Larson, Paul Richardson

Objective Electronic health records (EHRs) are a promising source of data for health outcomes research in oncology. A challenge in using EHR data is that selecting cohorts of patients often requires information in unstructured parts of the record. Machine learning has been used to address this, but even high-performing algorithms may select patients in a non-random manner and bias the resulting cohort. To improve the efficiency of cohort selection while measuring potential bias, we introduce a technique called Model-Assisted Cohort Selection (MACS) with Bias Analysis and apply it to the selection of metastatic breast cancer (mBC) patients. Materials and Methods We trained a model on 17,263 patients using term-frequency inverse-document-frequency (TF-IDF) and logistic regression. We used a test set of 17,292 patients to measure algorithm performance and perform Bias Analysis. We compared the cohort generated by MACS to the cohort that would have been generated without MACS as reference standard, first by comparing distributions of an extensive set of clinical and demographic variables and then by comparing the results of two analyses addressing existing example research questions. Results Our algorithm had an area under the curve (AUC) of 0.976, a sensitivity of 96.0%, and an abstraction efficiency gain of 77.9%. During Bias Analysis, we found no large differences in baseline characteristics and no differences in the example analyses. Conclusion MACS with bias analysis can significantly improve the efficiency of cohort selection on EHR data while instilling confidence that outcomes research performed on the resulting cohort will not be biased.

📄 PDF Abstract BibTeX arXiv:2001.09765

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Dictionary-based Pathology Mining with Hard-instance-assisted Classifier Debiasing for Genetic Biomarker Prediction from WSIs

2026-03-26 · Ling Zhang, Boxiang Yun, Ting Jin, Qingli Li 외 arxiv

Prediction of genetic biomarkers, e.g., microsatellite instability in colorectal cancer is crucial for clinical decision making. But, two primary challenges hamper accurate prediction: (1) It is difficult to construct a …

Decision Making

Characterization and reduction of variability in selection based on effect-size using association measures in cohort study of heterogeneous diseases

2017-05-27

Cohort studies employ pairwise measures of association to quantify dependencies among conditions and exposures. To reliably use these measures to draw conclusions about the underlying association strengths requires that …

Joint Selection: Adaptively Incorporating Public Information for Private Synthetic Data

2024-03-12 · Miguel Fuentes, Brett Mullins, Ryan McKenna, Gerome Miklau 외

Mechanisms for generating differentially private synthetic data based on marginals and graphical models have been successful in a wide range of settings. However, one limitation of these methods is their inability to inc…

Synthetic Data Generation

Multi-Cohort Framework with Cohort-Aware Attention and Adversarial Mutual-Information Minimization for Whole Slide Image Classification

2024-09-17 · Sharon Peled, Yosef E. Maruvka, Moti Freiman

Whole Slide Images (WSIs) are critical for various clinical applications, including histopathological analysis. However, current deep learning approaches in this field predominantly focus on individual tumor types, limit…

Diversityimage-classificationImage Classificationwhole slide images

Early completion based on multiple dosages to accelerate maximum tolerated dose-finding

2021-09-22 · Masahiro Kojima

Background: Phase I trials desire to identify the maximum tolerated dose (MTD) early and proceed quickly to an expansion cohort or phase II trial for efficacy. We propose an early completion method based on multiple dosa…