Agent harnesses need to make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System‑1 decision models answer these with a single forward pass that outputs class probabilities, promising large cost and latency savings.
We performed a paired evaluation of two System‑1 models—open‑weight Laya and hosted Jev—across 11 decision points built from 18 public sources. The benchmark contains 7,283 base cases and 6,640 robustness variants, all with byte‑identical inputs, and includes paired tests plus cross‑hardware and cross‑day reproducibility checks.
Jev outperformed Laya on 9 of 11 points, improving accuracy by 10.8 to 46.0 percentage points. Neither model beat chance on zero‑shot model routing, and they tied on RAG relevance gating. Laya’s answers changed 30% when option order was reversed and degraded sharply as candidate count or similarity grew (31% error with 50 nearest‑neighbour tools, versus 2% for Jev on items with a unique correct tool).
We also audited our own pipeline and uncovered three analysis errors and one design confound that distorted headline claims: an omitted pre‑screen cost (reported 23.9% saving, actual 4.3%), gate accuracy reported as end‑to‑end quality (58% vs 98%), in‑sample thresholds (5% target, up to 17% held‑out misses), and a “channel effect” on injection false positives that vanished with channel‑native content. Two other suspected confounds did not change the conclusions. All cases, raw outputs, and analysis code are available at https://github.com/David-DL-Space/sys1-eval.
Review