We introduce PAI‑Bench, a provider‑neutral benchmark designed to evaluate persistent AI agents on their fidelity to a versioned, update‑governed identity contract. The benchmark separates identity‑related capabilities into recall, composition, behavioral enactment, resistance, persistence, lineage, and role‑conditioned updates, while keeping scoring oracles outside the target process to avoid evaluator interference.
Two frozen campaigns were run, covering sixteen synthetic profiles, thirty‑two probes, and three independently initialized target configurations, yielding a total of 1,536 retained responses. An independent literal audit found that all 48 atomic responses contained direct‑parent identifiers, but only one included an implicit self‑portrait.
For eight profiles, adding explicit field cues raised the joint presence of three identity identifiers from 0/8 to 7/8 under the same four‑sentence instruction. A separate startup body‑label substitution experiment increased full‑designation presence from 1/8 to 7/8, while parent identifiers remained absent. These contrasts reveal prompt‑dependent component selection and component‑specific sensitivity to startup cues in the tested deployments.
Repeating identical factorial responses showed that Claude’s headline mean score was 12.5 percentage points lower than Astra’s, demonstrating that evaluator sensitivity can vary independently of target behavior. All studies used a single target sample per condition, followed by post‑hoc audits and follow‑up analyses.
PAI‑Bench provides a reproducible evaluation protocol for measuring factual availability, identity expression, and behavioral enactment as distinct aspects of identity‑contract fidelity.
Review