F.I.B. / Suites

Suites

Each suite is an independent evaluation with its own fixtures, scorer, and checked-in baseline. A run always belongs to exactly one suite; cross-suite scores are combined only through the locked weight vector.

presnapPre-snap reasoningdryliveFootball IQ · w 0.15

Read a formation before the snap: alignment, personnel, leverage, and the structural tells a coach checks in the first two seconds.

Keyed answers scored deterministically; the live path sends the identical prompt to every candidate model through the provider-fair adapter. 120-case candidate wave, pinned generation seed.

situationalSituational decisionsdryliveFootball IQ · w 0.15

Down, distance, clock, score: does the model make the call a competent staff would make, in context?

Keyed expected calls with acceptable-alternative sets. The current auditable source covers clock/timeout cases; the wider wave remains unverified until its digest is promoted.

football-qaFootball QA + ontologydryliveFootball IQ · w 0.10

Pure football knowledge: rules, schemes, coverages, terminology, and coaching decisions, graded by difficulty tier.

Keyed multiple-choice with a published gameability audit: null heuristics (always-longest, always-shortest, always-first) are computed per segment and gated within ±10pp of chance, so a football-blind test-taker cannot beat the suite.

arenaArena head-to-headliveFootball IQ · w 0.05

Two models call the same 12 seeded snaps against each other over a deterministic simulator, sides swapped, one match at a time.

Identical JSON-only decision prompt to both models with identical local validation; an invalid call forfeits the snap to a neutral default and is recorded. The underlying simulation-truth contract is provisional and every Arena surface is labeled accordingly.

ai-editAtomic editorliveHarness Fit · w 0.25

Exercise the production atomic play-edit route with 32 deterministic cases across four difficulty tiers.

Final-play changes keyed by player ID; collateral edits rejected; the expected tool path is pinned; results gated by both the play scorer and the play linter. The production route owns the model — no overrides.

concepts-canvasConcepts + canvasdryHarness Fit · w 0.20

Can the system draw the asked-for play on the canvas: concepts, routes, and alignments as structured, lintable artifacts.

Deterministic fixture expectations including deliberate negative mutations. Dry-only by design; the observed-capability baseline is tracked separately from fixture-contract resolution so known gaps are never laundered into the score.

second-brainSecond Brain retrievalunavailableHarness Fit · w 0.10

Retrieval quality over a coach’s own knowledge vault through the production chat and graph routes.

Unavailable until a benchmark-only, isolated seeded dataset exists. It carries weight in the vector but contributes nothing until then — the composite renormalizes and says so.

playgent-contractPlayGent contractliveUnweighted · w 0.00

End-to-end contract of the production coaching agent route: right answer AND right tool path, streamed over SSE.

Cases can pin the tool-call path, so a correct artifact produced through the wrong tools still fails. Deliberately outside the locked weight vector: present in every report as "present, not yet weighed."

“dry” = deterministic fixture-contract runs (regression evidence about fixtures and keys). “live” = real provider calls, spend-capped (model-quality evidence). The two evidence classes are labeled and never mixed.