paper-with-me

홈 › Papers

Analysis and classification of main risk factors causing stroke in Shanxi Province

2021-05-29 · Junjie Liu, Yiyang Sun, Jing Ma, Jiachen Tu, Yuhui Deng, Ping He, Huaxiong Huang, Xiaoshuang Zhou, Shixin Xu

In China, stroke is the first leading cause of death in recent years. It is a major cause of long-term physical and cognitive impairment, which bring great pressure on the National Public Health System. Evaluation of the risk of getting stroke is important for the prevention and treatment of stroke in China. A data set with 2000 hospitalized stroke patients in 2018 and 27583 residents during the year 2017 to 2020 is analyzed in this study. Due to data incompleteness, inconsistency, and non-structured formats, missing values in the raw data are filled with -1 as an abnormal class. With the cleaned features, three models on risk levels of getting stroke are built by using machine learning methods. The importance of "8+2" factors from China National Stroke Prevention Project (CSPP) is evaluated via decision tree and random forest models. Except for "8+2" factors the importance of features and SHAP1 values for lifestyle information, demographic information, and medical measurement are evaluated and ranked via a random forest model. Furthermore, a logistic regression model is applied to evaluate the probability of getting stroke for different risk levels. Based on the census data in both communities and hospitals from Shanxi Province, we investigate different risk factors of getting stroke and their ranking with interpretable machine learning models. The results show that Hypertension (Systolic blood pressure, Diastolic blood pressure), Physical Inactivity (Lack of sports), and Overweight (BMI) are ranked as the top three high-risk factors of getting stroke in Shanxi province. The probability of getting stroke for a person can also be predicted via our machine learning model.

📄 PDF Abstract BibTeX arXiv:2106.00002

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningInterpretable Machine LearningMissing Values

Methods 이 논문이 사용한 방법론

Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

Range contraction enables harvesting to extinction

2017-03-30

Economic incentives to harvest a species usually diminish as its abundance declines, because harvest costs increase. This prevents harvesting to extinction. A known exception can occur if consumer demand causes a declini…

Disentangling genetic and environmental risk factors for individual diseases from multiplex comorbidity networks

2016-05-31

Most disorders are caused by a combination of multiple genetic and/or environmental factors. If two diseases are caused by the same molecular mechanism, they tend to co-occur in patients. Here we provide a quantitative m…

ETF Risk Models

2021-10-14 · Zura Kakushadze, Willie Yu

We discuss how to build ETF risk models. Our approach anchors on i) first building a multilevel (non-)binary classification/taxonomy for ETFs, which is utilized in order to define the risk factors, and ii) then building …

Binary Classification

Covid-19 risk factors: Statistical learning from German healthcare claims data

2021-02-04 · Roland Jucknewitz, Oliver Weidinger, Anja Schramm

We analyse prior risk factors for severe, critical or fatal courses of Covid-19 based on a retrospective cohort using claims data of the AOK Bayern. As our main methodological contribution, we avoid prior grouping and pr…

Schrödinger Risk Diversification Portfolio

2022-02-21 · Yusuke Uchiyama, Kei Nakagawa

The mean-variance portfolio that considers the trade-off between expected return and risk has been widely used in the problem of asset allocation for multi-asset portfolios. However, since it is difficult to estimate the…

Dimensionality Reduction