Benchy is a semantic language and execution engine designed for benchmarking AI programs. A benchmark is fully specified by a program, a scoring function, and a dataset, denoted as $$B = (P, S, D)$$, and it is independent of the AI system that runs it. A single run binds the benchmark with an AI system, expressed as $$R = (B, AI)$$.
Benchmarks are authored in canonical YAML, where each semantic concept has a single valid syntax and is classified under a shared task/domain/language ontology. During compilation the YAML is deterministically transformed into a canonical JSON intermediate representation that the engine executes. Compilation only changes the representation, never repairing invalid definitions or injecting hidden defaults.
Programs follow fixed schemas of named input and output fields; the leaf output fields correspond directly to scoring dimensions. The engine exposes a universal runtime contract: it accepts a named‑field input object and returns a named‑field output object. External AI systems adapt only at this boundary, ensuring that integration mechanics never leak into benchmark semantics.
The paper details the following core aspects:
- The semantic object model and its attributes
- The ontology and the validation rules mapping tasks to programs
- Scoring and failure semantics
- The overall compilation and execution architecture
- The current language’s scope and limitations
An appendix formalizes the normative engineering contract for the first engine implementation, specifying implementation details and compatibility requirements.
Review: By providing a unified YAML‑>JSON compilation pipeline and a single runtime interface, Benchy cleanly separates benchmark definition from AI system implementation, offering a scalable and verifiable infrastructure for task‑oriented evaluation across diverse platforms and models.