Abstract
The increasing availability of longitudinal health data presents significant opportunities to improve disease risk prediction and support evidence-based decision-making. This study evaluates the application of supervised machine learning models to predict 10-year coronary heart disease (CHD) risk using data from the Framingham Heart Study. Given the substantial clinical and economic burden associated with cardiovascular disease, improving early risk identification has important implications for preventive healthcare planning and resource allocation.
Five machine learning algorithms—Logistic Regression, Random Forest, Gradient Boosting, K-Nearest Neighbours, and XGBoost—were developed and compared. The models integrate traditional clinical indicators and lifestyle-related variables to examine their combined predictive value. Performance was evaluated using accuracy, precision, recall, F1-score, and ROC-AUC, with stratified cross-validation applied to enhance robustness and SMOTE used to address class imbalance.
Logistic Regression achieved the highest overall accuracy and ROC-AUC, while ensemble methods demonstrated a more balanced relationship between precision and recall. However, recall remained modest across all models, highlighting the persistent challenge of detecting minority cardiovascular events in imbalanced datasets. Feature importance analysis consistently identified systolic blood pressure, age, BMI, cholesterol, and glucose levels as key predictors, aligning with established cardiovascular risk theory.
The findings contribute to applied predictive analytics by evaluating model trade-offs between accuracy and sensitivity in a high-stakes healthcare context. Beyond clinical relevance, the study demonstrates the managerial value of predictive modelling for healthcare planning, insurance risk assessment, and population health management.