F.I.B. / Methodology

Methodology

A benchmark run by the company whose product it measures has one honest path: publish the rules, the failure modes, and the null results, and make the gaming harder for ourselves than for anyone else. These are the rules the harness enforces structurally — not by convention.

Scoring

How a score is computed

FIB = Σ (suite_weight × suite_pass_rate), computed per model over F.I.B. Verified cases only. Two composites are reported — Football IQ (provider-fair prompts, identical to every model) and Harness Fit (Playi’s production routes) — and they are never blended into one number, because “understands football” and “works well inside our product” are different claims.

The weight vector was locked on 2026-08-06 from the measured production operation mix (what coaches actually run: chat, play generation, edits, QA). It is recorded with a date beside every published score so a number remains auditable later.

Live runs send each fixture’s prompt through a provider-fair adapter: identical prompt, identical local JSON parsing, no provider gets a structured-output crutch another lacks. Responses are scored by the suite’s own keyed scorer. Live spend is capped per run, and the cap is part of the published record.

Integrity

The honesty rules

Verified-only headline

The composite F.I.B. Score is computed over F.I.B. Verified cases only. A suite with zero verified cases is excluded and the remaining weights renormalize, with the renormalization flagged in the report. The score is null, never 0, when nothing is verified.

Sample size travels with every rate

Every pass rate is published with its n. Leaderboard intervals are Wilson score intervals, not normal approximations on tiny samples.

Provider errors are not model errors

A failed provider call is recorded as a distinct provider_error case. An outage is never laundered into "the model answered badly."

Partial runs are labeled

A run cut short by a spend cap or provider errors is labeled partial at the suite, composite, and model level. It is never silently averaged as if it were a complete sweep.

Gameability is audited, and published

Multiple-choice suites carry a model-free null-heuristic audit (always-longest, always-shortest, always-first-position) gated within ±10 percentage points of chance in either direction. The computed null baselines ship inside the suite baseline, so any reader can see what a football-blind heuristic scores.

Dry evidence is not live evidence

Deterministic fixture-contract runs prove fixture and key behavior — regression evidence, not model-quality evidence. The two evidence classes are labeled and never mixed.

Right answer AND right path

Agentic suites can pin the expected tool-call path. A correct artifact produced through the wrong tool path still fails the case.

Humans admit every fixture

Failures grow the eval set through a triage loop, but nothing lands in a fixture set except through a reviewed pull request. Verification promotion is an explicit human act recorded as a digest beside the wave.

Labels

Result labels

verified — the case wave passed human verification review; its exact content digest is promoted and recorded. Only verified cases feed the headline composites.

candidate — the wave is authored and running but unpromoted. Candidate tallies are published for the difficulty readout, always under an unverified banner, and never feed the headline number.

partial — the run hit a spend cap or provider errors before completing; labeled at suite, composite, and model level.

private / self-reported — anything we cannot reproduce from the harness ourselves is labeled as such, or not published.

Suite-by-suite scoring detail lives on the .