Generative Adversarial Networks for Synthetic Data Generation in Diabetic Patient Research: Techniques, Applications, and Challenges.

Synthetic data generation is a strategy used to address the lack and complex process to acquire clinical data and information, in particular in type 2 diabetes mellitus (T2DM) research. T2DM is characterized by chronic hyperglycemia with macrovascular and microvascular complications. Nevertheless, d...

Descripción completa

Detalles Bibliográficos
Publicado en:Studies in Health Technology & Informatics Vol. 330; pp. 714 - 750
Autores principales: GARCÍA-DOMÍNGUEZ, Antonio, ACOSTA-JIMÉNEZ, Samara, GONZALEZ-CURIEL, Irma, VILLAGRANA-BAÑUELOS, Karen E., ACOSTA-CRUZ, Erika, GALVÁN-TEJADA, Jorge I., GALVAN-TEJADA, Carlos E.
Formato: equations & formulas pictorial research tables/charts Journal Article
Publicado: Sage Publications Inc. 2025
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Synthetic data generation is a strategy used to address the lack and complex process to acquire clinical data and information, in particular in type 2 diabetes mellitus (T2DM) research. T2DM is characterized by chronic hyperglycemia with macrovascular and microvascular complications. Nevertheless, despite the importance of data to improve diagnostic accuracy, better treatments, and personalized patient care, medical datasets are often restricted by ethical and privacy constraints. In this sense, this chapter evaluates four synthetic data generation techniques, Gaussian Mixture Models (GMM), Generative Adversarial Networks (GAN), Wasserstein GAN (WGAN), and Variational Autoencoders (VAE). The quality of the generated data was assessed through statistical divergence metrics—specifically Jensen-Shannon (JSD) and Kullback-Leibler (KLD) and by analyzing their im-pact on classification performance. The results indicate that GMM achieved the lowest JSD, showing the best overall distributional similarity, while WGAN ob-tained the lowest KLD, suggesting a closer alignment in information content with real data. Additionally, GAN and WGAN demonstrated the highest predictive per-formance in classification tasks, indicating that they better preserved essential relationships within the data. These findings confirm that generative strategies of using synthetic data to im-prove T2DM research are feasible, offering an alternative to develop diagnosis tools without compromising patient confidentiality. It is possible to conclude that the generation method selection depends on the type of data and research objective, maximizing statistical similarity, optimizing performance, or balancing both aims. Synthetic data generation approaches represent a feasible approach to expand bal-anced and quality datasets to advance in personalized healthcare for diabetes patients.