Methodology
A benchmark run by the company whose product it measures has one honest path: publish the rules, the failure modes, and the null results, and make the gaming harder for ourselves than for anyone else. These are the rules the harness enforces structurally — not by convention.
How a score is computed
FIB = Σ (suite_weight × suite_pass_rate), computed per model over F.I.B. Verified cases only. Two composites are reported — Football IQ (provider-fair prompts, identical to every model) and Harness Fit (Playi’s production routes) — and they are never blended into one number, because “understands football” and “works well inside our product” are different claims.
The weight vector was locked on 2026-08-06 from the measured production operation mix (what coaches actually run: chat, play generation, edits, QA). It is recorded with a date beside every published score so a number remains auditable later.
Live runs send each fixture’s prompt through a provider-fair adapter: identical prompt, identical local JSON parsing, no provider gets a structured-output crutch another lacks. Responses are scored by the suite’s own keyed scorer. Live spend is capped per run, and the cap is part of the published record.
The honesty rules
The composite F.I.B. Score is computed over F.I.B. Verified cases only. A suite with zero verified cases is excluded and the remaining weights renormalize, with the renormalization flagged in the report. The score is null, never 0, when nothing is verified.
Every pass rate is published with its n. Leaderboard intervals are Wilson score intervals, not normal approximations on tiny samples.
A failed provider call is recorded as a distinct provider_error case. An outage is never laundered into "the model answered badly."
A run cut short by a spend cap or provider errors is labeled partial at the suite, composite, and model level. It is never silently averaged as if it were a complete sweep.
Multiple-choice suites carry a model-free null-heuristic audit (always-longest, always-shortest, always-first-position) gated within ±10 percentage points of chance in either direction. The computed null baselines ship inside the suite baseline, so any reader can see what a football-blind heuristic scores.
Deterministic fixture-contract runs prove fixture and key behavior — regression evidence, not model-quality evidence. The two evidence classes are labeled and never mixed.
Agentic suites can pin the expected tool-call path. A correct artifact produced through the wrong tool path still fails the case.
Failures grow the eval set through a triage loop, but nothing lands in a fixture set except through a reviewed pull request. Verification promotion is an explicit human act recorded as a digest beside the wave.
Result labels
verified — the case wave passed human verification review; its exact content digest is promoted and recorded. Only verified cases feed the headline composites.
candidate — the wave is authored and running but unpromoted. Candidate tallies are published for the difficulty readout, always under an unverified banner, and never feed the headline number.
partial — the run hit a spend cap or provider errors before completing; labeled at suite, composite, and model level.
private / self-reported — anything we cannot reproduce from the harness ourselves is labeled as such, or not published.
Suite-by-suite scoring detail lives on the .