NeFut Logo NeFut
Admin Login

[CS.AI] FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models

Published at: 2026-09-04 22:00 Last updated: 2026-09-05 12:23
#AI #Machine Learning #LLM

Evaluating large language models (LLMs) in safety‑critical, physics‑governed settings requires more than accuracy‑centric metrics. Predictions that are numerically close to the ground truth can still breach operational constraints, combine fields in physically inconsistent ways, or fail to produce usable structured outputs. Existing protocols do not reliably capture these failure modes.

We introduce FLY‑EVAL++, an evidence‑driven evaluation protocol that deterministically checks protocol compliance, physical feasibility, and safety constraints, then aggregates the results with a fixed rubric into interpretable multi‑dimensional scores.

We instantiate FLY‑EVAL++ for Flight Trajectory and Attitude Prediction (FTAP) by extending the PilotBench setting to include history‑conditioned and multi‑step prediction tasks. Across 66 LLMs, safety compliance emerges as the most discriminative behavior dimension: models with similar predictive performance differ by more than 28 points in safety score. Recurrent failures include safety violations under physically plausible predictions and instability in multi‑step rollouts.

These findings demonstrate that evaluation in safety‑critical domains should explicitly measure constraint satisfaction and structured validity rather than rely solely on accuracy‑centric reporting.

Review

Original Source: https://arxiv.org/abs/2609.04021

[h] Back to Home