Milia Habib, Marwa Bikai, Zaher Merhi, Rabih Rammal and Ali El Masri
Adv. Artif. Intell. Mach. Learn., - (-):-
1. Milia Habib: Department of Computer & Communications Engineering, Lebanese International University
2. Marwa Bikai: Department of Computer & Communications Engineering, Lebanese International University
3. Zaher Merhi: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon
4. Rabih Rammal: Department of Electrical Engineering Lebanese International University Beirut, Lebanon
5. Ali El Masri: Department of Computer & Communications Engineering Lebanese International University Beirut, Lebanon
DOI: 10.54364/AAIML.2026.65341
Article History: Received on: 23-May-26, Accepted on: 01-Sep-26, Published on: 07-Sep-26
Corresponding Author: Milia Habib
Email: milia.habib@liu.edu.lb
Citation: Marwa Bikai, et al. Calibration-Aware Hybrid Machine Learning and Deep Learning Ensemble for Cardiovascular Disease Prediction. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.65341
Cardiovascular disease (CVD) is a
major global cause of morbidity and mortality. This paper proposes a
calibration-aware hybrid ensemble framework. It combines traditional machine
learning (ML) and deep learning (DL) techniques to predict CVD from tabular
clinical datasets of 70,000 records, evaluated with leakage-safe nested
cross-validation and calibrated base probabilities, and reports both
discrimination and calibration metrics. Six machine learning models, including
LR, RF, ET, XGB, LGBM, and KNN, along with three deep learning models, namely
MLP, TabNet, and FT-Transformer, are evaluated. The top four ML models and the
three DL models are used to construct an ensemble via stacking and soft voting,
with both weighted and unweighted variants. The proposed approach integrates
hybrid ML/DL stacking, selective probability calibration, leakage-safe nested
cross-validation, and feature engineering. The performance is evaluated through
nine metrics (ROC-AUC, PR-AUC, accuracy, recall, precision, F1-score, Brier
score, log loss, and ECE). The results show that soft voting of XGB and
FT-Transformer achieved the highest discrimination performance with an ROC-AUC
of 80.16% and the best calibration metrics, including a Brier Score of 0.1803,
a Log Loss of 0.505, and ECE of 0.004. While the models demonstrated strong
internal performance, further validation on external datasets is needed to
establish their potential utility for clinical decision support.