Generative Adversarial Networks for Synthetic Data Generation in Diabetic Patient Research: Techniques, Applications, and Challenges.

Synthetic data generation is a strategy used to address the lack and complex process to acquire clinical data and information, in particular in type 2 diabetes mellitus (T2DM) research. T2DM is characterized by chronic hyperglycemia with macrovascular and microvascular complications. Nevertheless, d...

Descripción completa

Detalles Bibliográficos
Publicado en:Studies in Health Technology & Informatics Vol. 330; pp. 714 - 750
Autores principales: GARCÍA-DOMÍNGUEZ, Antonio, ACOSTA-JIMÉNEZ, Samara, GONZALEZ-CURIEL, Irma, VILLAGRANA-BAÑUELOS, Karen E., ACOSTA-CRUZ, Erika, GALVÁN-TEJADA, Jorge I., GALVAN-TEJADA, Carlos E.
Formato: equations & formulas pictorial research tables/charts Journal Article
Publicado: Sage Publications Inc. 2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=193074371&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 193074371
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        09269630
        U1V
      jtl: Studies in Health Technology & Informatics
      issn: 09269630
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 330
      pid: 344
      pub: Sage Publications Inc.
      place: Thousand Oaks, California
    artinfo:
      ui:
        193074371
        193074371
        193074371
        10.3233/SHTI251459
        193074371
      ppf: 714
      ppct: 36
      formats:
      tig:
        atl: Generative Adversarial Networks for Synthetic Data Generation in Diabetic Patient Research: Techniques, Applications, and Challenges.
      aug:
        au:
          GARCÍA-DOMÍNGUEZ, Antonio
          ACOSTA-JIMÉNEZ, Samara
          GONZALEZ-CURIEL, Irma
          VILLAGRANA-BAÑUELOS, Karen E.
          ACOSTA-CRUZ, Erika
          GALVÁN-TEJADA, Jorge I.
          GALVAN-TEJADA, Carlos E.
        affil: Unidad Académica de Ingeniería Eléctrica, Universidad Autónoma de Zacatecas, Zacatecas, Zacatecas, México.
      sug:
        subj:
          Generative Adversarial Networks
          Data Collection, Computer Assisted Methods
          Diabetes Mellitus, Type 2
          Research, Medical
          Models, Statistical
          Autoencoder
          Data Quality
          Prediction Models
          Human
          Diabetic Patients
          Male
          Female
          Adult
          Middle Age
          Comparative Studies
          Validation Studies
          Descriptive Statistics
          Boosting Machine Learning Algorithms
          Classification Algorithms
          Diabetes Mellitus, Type 2 Complications
          Diabetes Mellitus, Type 2 Physiopathology
          Biological Markers
          Conceptual Framework
          Diabetes Mellitus, Type 2 Therapy
          Diabetes Mellitus, Type 2 Diagnosis
          Individualized Medicine
          Diabetes Mellitus, Type 2 Etiology
          Diabetes Mellitus, Type 2 Risk Factors
          Risk Assessment
          Inflammation
          Wound Healing
          Adult: 19-44 years
          Middle Aged: 45-64 years
          Male
          Female
      ab: Synthetic data generation is a strategy used to address the lack and complex process to acquire clinical data and information, in particular in type 2 diabetes mellitus (T2DM) research. T2DM is characterized by chronic hyperglycemia with macrovascular and microvascular complications. Nevertheless, despite the importance of data to improve diagnostic accuracy, better treatments, and personalized patient care, medical datasets are often restricted by ethical and privacy constraints. In this sense, this chapter evaluates four synthetic data generation techniques, Gaussian Mixture Models (GMM), Generative Adversarial Networks (GAN), Wasserstein GAN (WGAN), and Variational Autoencoders (VAE). The quality of the generated data was assessed through statistical divergence metrics—specifically Jensen-Shannon (JSD) and Kullback-Leibler (KLD) and by analyzing their im-pact on classification performance. The results indicate that GMM achieved the lowest JSD, showing the best overall distributional similarity, while WGAN ob-tained the lowest KLD, suggesting a closer alignment in information content with real data. Additionally, GAN and WGAN demonstrated the highest predictive per-formance in classification tasks, indicating that they better preserved essential relationships within the data. These findings confirm that generative strategies of using synthetic data to im-prove T2DM research are feasible, offering an alternative to develop diagnosis tools without compromising patient confidentiality. It is possible to conclude that the generation method selection depends on the type of data and research objective, maximizing statistical similarity, optimizing performance, or balancing both aims. Synthetic data generation approaches represent a feasible approach to expand bal-anced and quality datasets to advance in personalized healthcare for diabetes patients.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N