Runtime failure monitors can exploit a model's internal representations to anticipate failures. This work audits that monitoring strategy on two autonomous‑driving tasks: online vectorized map generation with LaneSegNet and end‑to‑end planning with VAD. We find that frame‑level errors are predictable at inference time.
For LaneSegNet, a supervised latent probe achieves $AUROC=0.780$ for high Chamfer error, representing the first post‑hoc frame‑level failure monitor for online vectorized map generation. For VAD, a supervised planning‑latent probe reaches $AUROC=0.868$ on mean‑ADE failures.
Our audit further shows that strong failure prediction does not require internal access. Using only LaneSegNet's prediction outputs yields $AUROC=0.825$; for VAD, ego state, driving command, and the planner's predicted trajectory alone achieve $AUROC=0.924$ on the same mean‑ADE endpoint. Adding latent features to either baseline provides no statistically significant improvement. This observation holds across a broad suite of planning failure endpoints, even those whose labels depend on geometry unavailable to the non‑latent baseline.
Thus, predicting failure from an internal representation does not prove that the representation offers useful information beyond observable inputs and outputs. We propose an evaluation protocol to test the incremental value of latent access and release per‑frame failure endpoint labels.
Review