Ensemble Learning for Stroke Classification Based on Generative Adversarial Networks.

Background and Objective: Class imbalance in stroke datasets often results in increased misdiagnosis rates, biased risk factor analysis, and limited clinical utility of machine learning models. This study proposed a method combining Generative Adversarial Networks (GAN) with a hard voting ensemble t...

Descripción completa

Detalles Bibliográficos
Publicado en:Health & Technology Vol. 16; no. 3; pp. 527 - 538
Autores principales: Hong, Xiaorui, Wang, Ping, Bai, Jingchen, Zhu, Suling
Formato: algorithm equations & formulas research tables/charts Journal Article
Publicado: Springer Nature May2026
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=193278021&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 193278021
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        21907188
        BEWM
      jtl: Health & Technology
      issn: 21907188
      maglogo: N
    pubinfo:
      dt: May2026
      vid: 16
      iid: 3
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        193278021
        192536276
        193278021
        193278021
        10.1007/s12553-026-01062-1
        193278021
      ppf: 527
      ppct: 11
      formats:
      tig:
        atl: Ensemble Learning for Stroke Classification Based on Generative Adversarial Networks.
      aug:
        au:
          Hong, Xiaorui
          Wang, Ping
          Bai, Jingchen
          Zhu, Suling
        affil: https://ror.org/01mkqqe32 School of Public Health, Lanzhou University, 730000, Lanzhou, Gansu, China
      sug:
        subj:
          Stroke Risk Factors
          Risk Assessment
          Stroke Classification
          Ensemble Learning
          Generative Adversarial Networks
          Prediction Models
          Prediction Algorithms
          Classification Algorithms
          Sensitivity and Specificity Evaluation
          Human
          Questionnaires
          Logistic Regression
          Comparative Studies
          Random Forest
          Boosting Machine Learning Algorithms
          Odds Ratio
          Age Factors
          Depression
          Income
          Poverty
          Fatty Acids, Unsaturated
          Food Intake
          Adult
          Middle Age
          Aged
          Aged, 80 and Over
          Self Report
          Stroke Patients
          Descriptive Statistics
          Data Analysis Software
          Nonexperimental Studies
          Validation Studies
          United States
          Male
          Female
          Psychological Tests
          Adult: 19-44 years
          Middle Aged: 45-64 years
          Aged: 65+ years
          Aged, 80 & over
          Male
          Female
      ab: Background and Objective: Class imbalance in stroke datasets often results in increased misdiagnosis rates, biased risk factor analysis, and limited clinical utility of machine learning models. This study proposed a method combining Generative Adversarial Networks (GAN) with a hard voting ensemble to enhance the accuracy of stroke risk identification, thereby supporting improved stroke prevention and control. Methods: Data were from 4,238 participants in the NHANES database (2007–2018). A balanced subset resampling strategy combined with LASSO-logistic regression was employed to identify stroke risk factors. GAN was used to generate minority class samples to achieve dataset balance, and its performance was compared with traditional data balancing methods. Four base classifiers were constructed using selected features, and predictions were integrated via hard voting, including Logistic Regression (LR), Random Forest (RF), XGBoost, and Gradient Boosting Decision Tree (GBDT). Performance was assessed using sensitivity, F1 score, and G-mean. Results: Age (OR = 1.050, P < 0.001) and depression score (DPQ-9, OR = 1.065, P < 0.001) significantly increased stroke risk. Family income-to-poverty ratio (PIR, OR = 0.911, P = 0.026) and total polyunsaturated fatty acid intake (PUFA, OR = 0.988, P = 0.046) exhibited protective effects. GAN-based data balancing improved XGBoost's sensitivity from 0.018 to 0.915. The voting ensemble achieved an F1 score of 0.956 and specificity of 1.000. SHapley Additive exPlanations (SHAP) analysis confirmed the dominant role of these factors. Conclusion: This study integrated GAN-based data generation with a hard voting mechanism to effectively address class imbalance in stroke prediction. It significantly improved model stability and clinical identification capability, providing a reliable tool to screen high-risk populations and formulate targeted prevention strategies.
      pubtype: Academic Journal
      doctype:
        algorithm
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N