paper-with-me

Papers

A system for exploring big data: an iterative k-means searchlight for outlier detection on open health data

2023-04-05 · A. Ravishankar Rao, Daniel Clarke, Subrata Garai, Soumyabrata Dey

The interactive exploration of large and evolving datasets is challenging as relationships between underlying variables may not be fully understood. There may be hidden trends and patterns in the data that are worthy of further exploration and analysis. We present a system that methodically explores multiple combinations of variables using a searchlight technique and identifies outliers. An iterative k-means clustering algorithm is applied to features derived through a split-apply-combine paradigm used in the database literature. Outliers are identified as singleton or small clusters. This algorithm is swept across the dataset in a searchlight manner. The dimensions that contain outliers are combined in pairs with other dimensions using a susbset scan technique to gain further insight into the outliers. We illustrate this system by anaylzing open health care data released by New York State. We apply our iterative k-means searchlight followed by subset scanning. Several anomalous trends in the data are identified, including cost overruns at specific hospitals, and increases in diagnoses such as suicides. These constitute novel findings in the literature, and are of potential use to regulatory agencies, policy makers and concerned citizens.

📄 PDF Abstract BibTeX arXiv:2304.02189

Code (0)

등록된 구현이 없습니다.

Tasks

Outlier Detection

Methods 이 논문이 사용한 방법론

k-Means Clustering k-Means Clustering is a clustering algorithm that divides a training set into $k$ different clusters of examples that are near each other. It works by initializing $k$…

Similar Papers 제목 키워드 기반

PIKS: A Technique to Identify Actionable Trends for Policy-Makers Through Open Healthcare Data

2023-04-05 · A. Ravishankar Rao, Subrata Garai, Soumyabrata Dey, Hang Peng

With calls for increasing transparency, governments are releasing greater amounts of data in multiple domains including finance, education and healthcare. The efficient exploratory analysis of healthcare data constitutes…

Outlier Detection

A Searchlight Factor Model Approach for Locating Shared Information in Multi-Subject fMRI Analysis

2016-09-29 · Hejia Zhang, Po-Hsuan Chen, Janice Chen, Xia Zhu 외

There is a growing interest in joint multi-subject fMRI analysis. The challenge of such analysis comes from inherent anatomical and functional variability across subjects. One approach to resolving this is a shared respo…

General Classification

A Convolutional Autoencoder for Multi-Subject fMRI Data Aggregation

2016-08-17 · Po-Hsuan Chen, Xia Zhu, Hejia Zhang, Javier S. Turek 외

Finding the most effective way to aggregate multi-subject fMRI data is a long-standing and challenging problem. It is of increasing interest in contemporary fMRI studies of human cognition due to the scarcity of data per…

Anatomy

Gradient-based Representational Similarity Analysis with Searchlight for Analyzing fMRI Data

2018-09-12 · Xiaoliang Sheng, Muhammad Yousefnezhad, Tonglin Xu, Ning Yuan 외

Representational Similarity Analysis (RSA) aims to explore similarities between neural activities of different stimuli. Classical RSA techniques employ the inverse of the covariance matrix to explore a linear model betwe…

Data Integration with Fusion Searchlight: Classifying Brain States from Resting-state fMRI

2024-12-13 · Simon Wein, Marco Riebel, Lisa-Marie Brunner, Caroline Nothdurfter 외

Resting-state fMRI captures spontaneous neural activity characterized by complex spatiotemporal dynamics. Various metrics, such as local and global brain connectivity and low-frequency amplitude fluctuations, quantify di…

Data IntegrationSpecificity