Interpretable heart disease risk prediction via FCA-constrained logistic regression.

Objective: To develop an interpretable and clinically coherent heart disease risk prediction model by integrating Formal Concept Analysis (FCA) with a novel closure-constrained logistic regression that enforces coefficient coherence within FCA-derived concepts. Methods: We used the Heart Disease Hea...

Descripción completa

Detalles Bibliográficos
Publicado en:Health Informatics Journal Vol. 32; no. 2; pp. 1 - 15
Autores principales: Salehi, Arman, Heydarian, Ashkan, Goudarzi, Hamid Reza, Rad, Zahra Farzin
Formato: equations & formulas research tables/charts Journal Article
Publicado: Sage Publications Inc. Apr-Jun2026
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Objective: To develop an interpretable and clinically coherent heart disease risk prediction model by integrating Formal Concept Analysis (FCA) with a novel closure-constrained logistic regression that enforces coefficient coherence within FCA-derived concepts. Methods: We used the Heart Disease Health Indicators dataset (BRFSS 2015; N≈380,000). Predictors were discretized into binary attributes, and closed itemsets were extracted via FCA. A closure penalty, which minimizes within-concept coefficient variance, was added to the logistic regression objective. Hyperparameters (closure strength λ, FCA minimum support) were selected using five-fold cross-validation on the training set. Baselines included L2-regularized logistic regression, Random Forest, and Gradient Boosting. Performance was evaluated on a held-out test set using AUC, accuracy, precision, recall, F1, PR-AUC, and Brier Score. Results: On the held-out test set, the FCA-constrained model achieved Accuracy = 0.906, AUC = 0.810, Precision = 0.709, Recall = 0.544, F1 = 0.556, PR-AUC = 0.265, and Brier Score = 0.078. Compared to baselines, the FCA model produced well-calibrated probabilities and the highest precision and F1-score, while providing concept-level explanations grounded in clinically coherent closed itemsets. Conclusion: Embedding FCA structure directly into model training yields an interpretable linear model with competitive discrimination and improved precision at clinically relevant thresholds.