This work introduces a reliability‑aware Hybrid‑K ensemble selection framework for multiclass cervical cytology classification, evaluated on the SIPaKMeD dataset. Nine deep learning architectures were benchmarked using a fixed stratified five‑fold split and three random seeds. After applying post‑hoc temperature scaling, models were measured with macro‑F1, accuracy, AUROC, expected calibration error (ECE), worst‑class ECE (WC‑ECE), area under the risk‑coverage curve (AURC), Brier score, and negative log‑likelihood (NLL). An equal‑weight composite score ranked the models, and the top performers were combined via soft voting into Hybrid‑K ensembles.
Robustness was examined through 5,000 Dirichlet‑sampled metric‑weight vectors, leave‑one‑metric‑out analysis, and corrected paired testing across 15 fold‑by‑seed evaluations. The final Hybrid‑2 ensemble, consisting of Swin‑Tiny and TinyViT‑5M, reduced AURC by 43%, NLL by 17%, and WC‑ECE by 36% relative to the best single model. It was selected in 96.8% of random weighting scenarios, remained unchanged across all leave‑one‑metric‑out analyses, and improved the overall composite score. Nevertheless, per‑metric improvements were not statistically significant after Holm‑Bonferroni correction (adjusted p = 0.168). Because post‑hoc calibration did not employ a fully independent calibration set, calibration‑dependent results should be regarded as exploratory internal estimates.
In summary, the proposed framework identified a compact ensemble that is robust to alternative metric weightings and enhances reliability point estimates under internal validation on a single dataset.
Review