Feature selection and risk prediction for patients with coronary artery disease using data mining.
Coronary artery disease (CAD) is an important cause of mortality across the globe. Early risk prediction of CAD would be able to reduce the death rate by allowing early and targeted treatments. In healthcare, some studies applied data mining techniques and machine learning algorithms on the risk pre...
| Publicado en: | Medical & Biological Engineering & Computing Vol. 58; no. 12; pp. 3123 - 3141 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Journal Article |
| Publicado: |
Springer Nature
2020
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=147105107&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 147105107 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 01400118 PO0 jtl: Medical & Biological Engineering & Computing issn: 01400118 maglogo: N pubinfo: dt: 2020 vid: 58 iid: 12 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 147105107 146844741 147105107 NLM33155096 10.1007/s11517-020-02268-9 NLM33155096 147105107 ppf: 3123 ppct: 18 formats: fmt: – @attributes: type: T – @attributes: type: P tig: atl: Feature selection and risk prediction for patients with coronary artery disease using data mining. aug: au: Md Idris, Nashreen Chiam, Yin Kia Varathan, Kasturi Dewi Wan Ahmad, Wan Azman Chee, Kok Han Liew, Yih Miin affil: Department of Software Engineering, Faculty of Computer Science and Information Technology, Universiti Malaya, 50603, Kuala Lumpur, Malaysia sug: subj: Coronary Arteriosclerosis Epidemiology Data Mining Algorithms Resource Databases ab: Coronary artery disease (CAD) is an important cause of mortality across the globe. Early risk prediction of CAD would be able to reduce the death rate by allowing early and targeted treatments. In healthcare, some studies applied data mining techniques and machine learning algorithms on the risk prediction of CAD using patient data collected by hospitals and medical centers. However, most of these studies used all the attributes in the datasets which might reduce the performance of prediction models due to data redundancy. The objective of this research is to identify significant features to build models for predicting the risk level of patients with CAD. In this research, significant features were selected using three methods (i.e., Chi-squared test, recursive feature elimination, and Embedded Decision Tree). Synthetic Minority Over-sampling Technique (SMOTE) oversampling technique was implemented to address the imbalanced dataset issue. The prediction models were built based on the identified significant features and eight machine learning algorithms, utilizing Acute Coronary Syndrome (ACS) datasets provided by National Cardiovascular Disease Database (NCVD) Malaysia. The prediction models were evaluated and compared using six performance evaluation metrics, and the top-performing models have achieved AUC more than 90%. Graphical abstract. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|