Early Prediction of Cardiovascular Disease Using Clinical Data
Keywords:
Cardiovascular Disease, Early Prediction, Clinical Data, Machine Learning, Gradient BoostingAbstract
Cardiovascular disease can present itself in an insidious manner, and detection of risk factors is crucial for preventing and managing the disease at an early stage. The study assessed the potential of commonly used clinical and behavioural parameters for predicting cardiovascular disease with five classification methods: logistic regression, decision tree, random forest, support vector machine and gradient boosting. Data comprised of demographic, physiological, laboratory and lifestyle data, and model performance was measured by accuracy, sensitivity, specificity, precision, F1 score, ROC–AUC and calibration. Gradient boosting performed best overall with a score of 86.8%, a sensitivity of 88.4%, a specificity of 85.4%, an F1 score of 86.3% and an ROC–AUC of 0.92. The most important predictors were systolic blood pressure, age, cholesterol, body mass index, fasting glucose, smoking status and family history. The results show that ensemble learning may identify the complex relationships existing between cardiovascular risk factors, and may help identify individuals at high risk earlier. The model provides supportive information for preventive screening, clinical prioritization and targeted lifestyle guidance in conjunction with clinical judgment. There is a need for external validation, fairness assessment and prospective testing prior to routine clinical use. Further studies also need to validate it against existing risk scores and establish if decisions based on the model have a positive impact on patient outcomes in various healthcare environments globally.
