From Text to Topics in Healthcare Records: An Unsupervised Graph Partitioning Methodology
Electronic Healthcare Records contain large volumes of unstructured data, including extensive free text. Yet this source of detailed information often remains under-used because of a lack of methodologies to extract interpretable content in a timely manner. Here we apply network-theoretical tools to analyse free text in Hospital Patient Incident reports from the National Health Service, to find clusters of documents with similar content in an unsupervised manner at different levels of resolution. We combine deep neural network paragraph vector text-embedding with multiscale Markov Stability community detection applied to a sparsified similarity graph of document vectors, and showcase the approach on incident reports from Imperial College Healthcare NHS Trust, London. The multiscale community structure reveals different levels of meaning in the topics of the dataset, as shown by descriptive terms extracted from the clusters of records. We also compare a posteriori against hand-coded categories assigned by healthcare personnel, and show that our approach outperforms LDA-based models. Our content clusters exhibit good correspondence with two levels of hand-coded categories, yet they also provide further medical detail in certain areas and reveal complementary descriptors of incidents beyond the external classification taxonomy.
Code (0)
등록된 구현이 없습니다.
Tasks
Community DetectionDescriptivegraph partitioningSimilar Papers 제목 키워드 기반
From Free Text to Clusters of Content in Health Records: An Unsupervised Graph Partitioning Approach
Electronic Healthcare records contain large volumes of unstructured data in different forms. Free text constitutes a large portion of such data, yet this source of richly detailed information often remains under-used in …
Community DetectionDescriptivegraph partitioningExtracting information from free text through unsupervised graph-based clustering: an application to patient incident records
The large volume of text in electronic healthcare records often remains underused due to a lack of methodologies to extract interpretable content. Here we present an unsupervised framework for the analysis of free text t…
ClusteringCommunity DetectionNeural Natural Language Processing for Unstructured Data in Electronic Health Records: a Review
Electronic health records (EHRs), digital collections of patient healthcare events and observations, are ubiquitous in medicine and critical to healthcare delivery, operations, and research. Despite this central role, EH…
Knowledge GraphsQuestion AnsweringWord EmbeddingsAn Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare Records
Electronic healthcare records are vital for patient safety as they document conditions, plans, and procedures in both free text and medical codes. Language models have significantly enhanced the processing of such record…
Adversarial RobustnessExplainable Artificial Intelligence (XAI)Feature ImportanceMedical Code PredictionContextual Embedding-based Clustering to Identify Topics for Healthcare Service Improvement
Understanding patient feedback is crucial for improving healthcare services, yet analyzing unlabeled short-text feedback presents significant challenges due to limited data and domain-specific nuances. Traditional superv…
Clustering