paper-with-me

Papers

The NLP Sandbox: an efficient model-to-data system to enable federated and unbiased evaluation of clinical NLP models

2022-06-28 · Yao Yan, Thomas Yu, Kathleen Muenzen, Sijia Liu, Connor Boyle, George Koslowski, Jiaxin Zheng, Nicholas Dobbins, Clement Essien, Hongfang Liu, Larsson Omberg, Meliha Yestigen, Bradley Taylor, James A Eddy, Justin Guinney, Sean Mooney, Thomas Schaffter

Objective The evaluation of natural language processing (NLP) models for clinical text de-identification relies on the availability of clinical notes, which is often restricted due to privacy concerns. The NLP Sandbox is an approach for alleviating the lack of data and evaluation frameworks for NLP models by adopting a federated, model-to-data approach. This enables unbiased federated model evaluation without the need for sharing sensitive data from multiple institutions. Materials and Methods We leveraged the Synapse collaborative framework, containerization software, and OpenAPI generator to build the NLP Sandbox (nlpsandbox.io). We evaluated two state-of-the-art NLP de-identification focused annotation models, Philter and NeuroNER, using data from three institutions. We further validated model performance using data from an external validation site. Results We demonstrated the usefulness of the NLP Sandbox through de-identification clinical model evaluation. The external developer was able to incorporate their model into the NLP Sandbox template and provide user experience feedback. Discussion We demonstrated the feasibility of using the NLP Sandbox to conduct a multi-site evaluation of clinical text de-identification models without the sharing of data. Standardized model and data schemas enable smooth model transfer and implementation. To generalize the NLP Sandbox, work is required on the part of data owners and model developers to develop suitable and standardized schemas and to adapt their data or model to fit the schemas. Conclusions The NLP Sandbox lowers the barrier to utilizing clinical data for NLP model evaluation and facilitates federated, multi-site, unbiased evaluation of NLP models.

📄 PDF Abstract BibTeX arXiv:2206.14181

Code (0)

등록된 구현이 없습니다.

Tasks

De-identification

Similar Papers 제목 키워드 기반

Digital Agriculture Sandbox for Collaborative Research

2025-11-20 · Osama Zafar, Rosemarie Santa González, Alfonso Morales, Erman Ayday arxiv

Digital agriculture is transforming the way we grow food by utilizing technology to make farming more efficient, sustainable, and productive. This modern approach to agriculture generates a wealth of valuable data that c…

Federated Learning

Federated Learning Playground

2026-02-23 · Bryan Shan, Alysa Ziying Tan, Han Yu arxiv

We present Federated Learning Playground, an interactive browser-based platform inspired by and extends TensorFlow Playground that teaches core Federated Learning (FL) concepts. Users can experiment with heterogeneous cl…

Federated Learning

FLoBC: A Decentralized Blockchain-Based Federated Learning Framework

2021-12-22 · Mohamed Ghanem, Fadi Dawoud, Habiba Gamal, Eslam Soliman 외

The rapid expansion of data worldwide invites the need for more distributed solutions in order to apply machine learning on a much wider scale. The resultant distributed learning systems can have various degrees of centr…

BIG-bench Machine LearningFederated Learning

DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models

2025-09-11 · Honghui Xu, Shiva Shrestha, Wei Chen, Zhiyuan Li 외 arxiv

As on-device large language model (LLM) systems become increasingly prevalent, federated fine-tuning enables advanced language understanding and generation directly on edge devices; however, it also involves processing s…

Federated Learning

Federated Causal Inference in Healthcare: Methods, Challenges, and Applications

2025-05-04 · Haoyang Li, Jie Xu, Kyra Gan, Fei Wang 외

Federated causal inference enables multi-site treatment effect estimation without sharing individual-level data, offering a privacy-preserving solution for real-world evidence generation. However, data heterogeneity acro…

Causal InferencePrivacy Preserving