Suites
Each suite is an independent evaluation with its own fixtures, scorer, and checked-in baseline. A run always belongs to exactly one suite; cross-suite scores are combined only through the locked weight vector.
Read a formation before the snap: alignment, personnel, leverage, and the structural tells a coach checks in the first two seconds.
Keyed answers scored deterministically; the live path sends the identical prompt to every candidate model through the provider-fair adapter. 120-case candidate wave, pinned generation seed.
Down, distance, clock, score: does the model make the call a competent staff would make, in context?
Keyed expected calls with acceptable-alternative sets. The current auditable source covers clock/timeout cases; the wider wave remains unverified until its digest is promoted.
Pure football knowledge: rules, schemes, coverages, terminology, and coaching decisions, graded by difficulty tier.
Keyed multiple-choice with a published gameability audit: null heuristics (always-longest, always-shortest, always-first) are computed per segment and gated within ±10pp of chance, so a football-blind test-taker cannot beat the suite.
Two models call the same 12 seeded snaps against each other over a deterministic simulator, sides swapped, one match at a time.
Identical JSON-only decision prompt to both models with identical local validation; an invalid call forfeits the snap to a neutral default and is recorded. The underlying simulation-truth contract is provisional and every Arena surface is labeled accordingly.
Exercise the production atomic play-edit route with 32 deterministic cases across four difficulty tiers.
Final-play changes keyed by player ID; collateral edits rejected; the expected tool path is pinned; results gated by both the play scorer and the play linter. The production route owns the model — no overrides.
Can the system draw the asked-for play on the canvas: concepts, routes, and alignments as structured, lintable artifacts.
Deterministic fixture expectations including deliberate negative mutations. Dry-only by design; the observed-capability baseline is tracked separately from fixture-contract resolution so known gaps are never laundered into the score.
Retrieval quality over a coach’s own knowledge vault through the production chat and graph routes.
Unavailable until a benchmark-only, isolated seeded dataset exists. It carries weight in the vector but contributes nothing until then — the composite renormalizes and says so.
End-to-end contract of the production coaching agent route: right answer AND right tool path, streamed over SSE.
Cases can pin the tool-call path, so a correct artifact produced through the wrong tools still fails. Deliberately outside the locked weight vector: present in every report as "present, not yet weighed."
“dry” = deterministic fixture-contract runs (regression evidence about fixtures and keys). “live” = real provider calls, spend-capped (model-quality evidence). The two evidence classes are labeled and never mixed.