ISSN :2582-9793

Calibration-Aware Hybrid Machine Learning and Deep Learning Ensemble for Cardiovascular Disease Prediction

Original Research (Published On: 07-Sep-2026 )
DOI : https://doi.org/10.54364/AAIML.2026.65341

Milia Habib, Marwa Bikai, Zaher Merhi, Rabih Rammal and Ali El Masri

Adv. Artif. Intell. Mach. Learn., - (-):-

1. Milia Habib: Department of Computer & Communications Engineering, Lebanese International University

2. Marwa Bikai: Department of Computer & Communications Engineering, Lebanese International University

3. Zaher Merhi: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon

4. Rabih Rammal: Department of Electrical Engineering Lebanese International University Beirut, Lebanon

5. Ali El Masri: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon

Download PDF Here

DOI: 10.54364/AAIML.2026.65341

Article History: Received on: 23-May-26, Accepted on: 01-Sep-26, Published on: 07-Sep-26

Corresponding Author: Milia Habib

Email: milia.habib@liu.edu.lb

Citation: Marwa Bikai, et al. Calibration-Aware Hybrid Machine Learning and Deep Learning Ensemble for Cardiovascular Disease Prediction. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.65341


Abstract

    

Cardiovascular disease (CVD) is a major global cause of morbidity and mortality. This paper proposes a calibration-aware hybrid ensemble framework. It combines traditional machine learning (ML) and deep learning (DL) techniques to predict CVD from tabular clinical datasets of 70,000 records, evaluated with leakage-safe nested cross-validation and calibrated base probabilities, and reports both discrimination and calibration metrics. Six machine learning models, including LR, RF, ET, XGB, LGBM, and KNN, along with three deep learning models, namely MLP, TabNet, and FT-Transformer, are evaluated. The top four ML models and the three DL models are used to construct an ensemble via stacking and soft voting, with both weighted and unweighted variants. The proposed approach integrates hybrid ML/DL stacking, selective probability calibration, leakage-safe nested cross-validation, and feature engineering. The performance is evaluated through nine metrics (ROC-AUC, PR-AUC, accuracy, recall, precision, F1-score, Brier score, log loss, and ECE). The results show that soft voting of XGB and FT-Transformer achieved the highest discrimination performance with an ROC-AUC of 80.16% and the best calibration metrics, including a Brier Score of 0.1803, a Log Loss of 0.505, and ECE of 0.004. While the models demonstrated strong internal performance, further validation on external datasets is needed to establish their potential utility for clinical decision support.

Statistics

   Article View: 87
   PDF Downloaded: 11