Activity‑cliff ranking remains challenging because tiny local structural changes can cause large activity differences, while high‑quality data that reveal the underlying mechanisms are scarce. To make better use of available activity labels, CliffRank combines absolute‑activity regression with ranking‑consistency learning. The framework trains two parallel predictors with a loss composed of mean squared error, a thresholded listwise loss, and Pairwise Preference Consistency (PPC), which aligns relative ordering in the preference‑probability space.
On three antimicrobial peptide datasets, CliffRank equipped with ESM2‑t12 achieved the highest mean Spearman correlation of 0.5393 and a mean Recall@50 of 21.4, although the leading method varied across individual datasets.
On three small‑molecule datasets, CliffRank paired with PNA and activating PPC after 120 epochs obtained the highest mean Spearman correlation of 0.6890, while its mean Recall@50 of 30.4 matched that of ACANet‑PNA. The PPC experiments also delineate its practical limits.
Asymmetric initialization improved the overall averages of MolCLR‑GIN but did not benefit every target. For PNA without pretrained weights, delayed PPC improved selected metrics, yet no schedule proved optimal for both mean Spearman correlation and mean Recall@50.
Future work should evaluate more targets and antimicrobial peptide systems, develop adaptive PPC schedules, and incorporate protein or membrane context when available.
Review