F.I.B.

The Football Intelligence Benchmark

F.I.B. measures whether an AI system actually understands football — not whether it can sound like it does. It is to agentic coaching what SWE-bench is to agentic coding: keyed answers, deterministic scorers, pinned tool paths, and production routes, instead of vibes-graded free text.

Scoring

Two composites, never blended

Football IQ
Which model understands football best?

Suites where every candidate model receives the identical prompt through a provider-fair adapter — identical parsing, no provider gets a structured-output crutch the other lacks.

Pre-snap reasoning0.15
Situational decisions0.15
Football QA + ontology0.10
Arena head-to-head0.05
Harness Fit
Which model works best inside Playi?

Suites that run through Playi’s own production routes, where the route may own or pin the model. This measures the system a coach actually touches, not a lab-only prompt.

Atomic editor0.25
Concepts + canvas0.20
Second Brain retrieval0.10

Weight vector locked 2026-08-06, derived from Playi’s measured production operation mix. Changing a weight requires an explicit sign-off recorded in the scoring module’s history. Suites outside the vector (the PlayGent route contract) appear in every report as “present, not yet weighed” rather than silently missing.

Current stateas of 2026-08-20

Honestly empty, on purpose

Every suite’s promotion digest is still unset: no wave has passed human verification review yet. The composite score is null — not zero — for every model, on both composites. Candidate-wave tallies exist internally and are always labeled as unverified; they never feed the headline number.

That is a deliberate design choice, enforced structurally in the scoring module: the headline number stays honestly empty until a human promotes a wave. Benchmarks have enough theater already.