paper-with-me

홈 › Papers

Automatically Labeling $200B Life-Saving Datasets: A Large Clinical Trial Outcome Benchmark

2024-06-13 · Chufan Gao, Jathurshan Pradeepkumar, Trisha Das, Shivashankar Thati, Jimeng Sun

Background: The global cost of drug discovery and development exceeds $200 billion annually, with clinical trial outcomes playing a critical role in the regulatory approval of new drugs and impacting patient outcomes. Despite their significance, large-scale, high-quality clinical trial outcome data are not readily available to the public, limiting advances in trial outcome predictive modeling. Methods: We introduce the Clinical Trial Outcome (CTO) knowledge base, a fully reproducible, large-scale (around 125K drug and biologics trials), open-source of clinical trial information including large language model (LLM) interpretations of publications, matched trials over phases, sentiment analysis from news, stock prices of trial sponsors, and other trial-related metrics. From this knowledge base, we additionally performed manual annotation of a set of recent clinical trials from 2020-2024. Results: We evaluated the quality of our knowledge base by generating high-quality trial outcome labels that demonstrate strong agreement with previously published expert annotations, achieving an F1 score of 94 for Phase 3 trials and 91 across all phases. Additionally, we benchmarked a suite of standard machine learning models on our manually annotated set, highlighting the distribution shift of recent trials and the need for continuously updated labeling methods. Conclusions: By analyzing CTO's performance on recent trials, we showed a need for recent, high-quality trial outcome labels. We release our knowledge base and labels to the public at https://chufangao.github.io/CTOD, which will also be regularly updated to support ongoing research in clinical trial outcomes, offering insights that could optimize the drug development process.

📄 PDF Abstract BibTeX arXiv:2406.10292

Code (0)

등록된 구현이 없습니다.

Tasks

Drug DiscoveryLarge Language ModelSentiment Analysis

Methods 이 논문이 사용한 방법론

BASE 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Promises and Pitfalls of Threshold-based Auto-labeling

2022-11-22 · NeurIPS 2023 11 · Harit Vishwakarma, Heguang Lin, Frederic Sala, Ramya Korlakai Vinayak

Creating large-scale high-quality labeled datasets is a major bottleneck in supervised machine learning workflows. Threshold-based auto-labeling (TBAL), where validation data obtained from humans is used to find a confid…

A new data augmentation method for intent classification enhancement and its application on spoken conversation datasets

2022-02-21 · Zvi Kons, Aharon Satt, Hong-Kwang Kuo, Samuel Thomas 외

Intent classifiers are vital to the successful operation of virtual agent systems. This is especially so in voice activated systems where the data can be noisy with many ambiguous directions for user intents. Before oper…

Active LearningData Augmentationintent-classificationIntent Classification

Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning

2017-03-16 · Mohammed Sadegh Norouzzadeh, Anh Nguyen, Margaret Kosmala, Ali Swanson 외

Having accurate, detailed, and up-to-date information about the location and behavior of animals in the wild would revolutionize our ability to study and conserve ecosystems. We investigate the ability to automatically, …

ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-Thought

2023-10-26 · Hanchong Zhang, Ruisheng Cao, Lu Chen, Hongshen Xu 외

Recently Large Language Models (LLMs) have been proven to have strong abilities in various domains and tasks. We study the problem of prompt designing in the text-to-SQL task and attempt to improve the LLMs' reasoning ab…

In-Context LearningText to SQLText-To-SQL

Automatic Pixelwise Object Labeling for Aerial Imagery Using Stacked U-Nets

2018-03-13 · Andrew Khalel, Motaz El-Saban

Automation of objects labeling in aerial imagery is a computer vision task with numerous practical applications. Fields like energy exploration require an automated method to process a continuous stream of imagery on a d…