A hybrid sampling algorithm combining M-SMOTE and ENN based on Random forest for medical imbalanced data.
The problem of imbalanced data classification often exists in medical diagnosis. Traditional classification algorithms usually assume that the number of samples in each class is similar and their misclassification cost during training is equal. However, the misclassification cost of patient samples...
| Publicado en: | Journal of Biomedical Informatics Vol. 107 |
|---|---|
| Autores principales: | , , , |
| Formato: | research Journal Article |
| Publicado: |
Academic Press Inc.
Jul2020
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=144627583&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 144627583 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: Jul2020 vid: 107 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 144627583 144627583 NLM32512209 144627583 10.1016/j.jbi.2020.103465 NLM32512209 144627583 ppct: 1 formats: tig: atl: A hybrid sampling algorithm combining M-SMOTE and ENN based on Random forest for medical imbalanced data. aug: au: Xu, Zhaozhao Shen, Derong Nie, Tiezheng Kou, Yue affil: School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China sug: subj: Algorithms Study Design Human Comparative Studies Multicenter Studies Evaluation Research Validation Studies ab: The problem of imbalanced data classification often exists in medical diagnosis. Traditional classification algorithms usually assume that the number of samples in each class is similar and their misclassification cost during training is equal. However, the misclassification cost of patient samples is higher than that of healthy person samples. Therefore, how to increase the identification of patients without affecting the classification of healthy individuals is an urgent problem. In order to solve the problem of imbalanced data classification in medical diagnosis, we propose a hybrid sampling algorithm called RFMSE, which combines the Misclassification-oriented Synthetic minority over-sampling technique (M-SMOTE) and Edited nearset neighbor (ENN) based on Random forest (RF). The algorithm is mainly composed of three parts. First, M-SMOTE is used to increase the number of samples in the minority class, while the over-sampling rate of M-SMOTE is the misclassification rate of RF. Then, ENN is used to remove the noise ones from the majority samples. Finally, RF is used to perform classification prediction for the samples after hybrid sampling, and the stopping criterion for iterations is determined according to the changes of the classification index (i.e. Matthews Correlation Coefficient (MCC)). When the value of MCC continuously drops, the process of iterations will be stopped. Extensive experiments conducted on ten UCI datasets demonstrate that RFMSE can effectively solve the problem of imbalanced data classification. Compared with traditional algorithms, our method can improve F-value and MCC more effectively. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|