F.I.B. / Leaderboard

Leaderboard

Per-model composite scores over F.I.B. Verified cases, with per-suite pass rates, sample sizes, Wilson confidence intervals, cost, and latency. Published runs are labeled verified, candidate, partial, or private — and a partial run is never averaged as if it were a complete sweep.

no verified public results yet

Every suite’s promotion digest is still unset: no wave has passed human verification review yet. The composite score is null — not zero — for every model, on both composites. Candidate-wave tallies exist internally and are always labeled as unverified; they never feed the headline number.

Internal candidate-wave runs exist across the suites, always labeled unverified. The first public rows land here once a fixture wave passes verification review and a multi-model live sweep is published against it — with its n, intervals, cost, and failure modes attached.

Football IQscore: null · verified n: 0

Which model understands football best? No model has a verified score yet; null is reported rather than a misleading zero.

Harness Fitscore: null · verified n: 0

Which model works best inside Playi? No model has a verified score yet; null is reported rather than a misleading zero.