Playi Labs

The science behind the agent.

Research, benchmarks, and infrastructure for football intelligence. Playi builds the AI agent football coaches actually use — Labs is where we publish how football understanding is measured, which models hold up, and the open infrastructure that makes the results reproducible.

Flagshipcandidate waves — no verified results yet

F.I.B. — Football Intelligence Benchmark

A rigorous evaluation of how well AI systems understand, reason about, and execute football concepts — 8 suites, two composites, deterministic scorers, and a verification bar no result crosses without human review.

Football IQscore: null
Which model understands football best?

Suites where every candidate model receives the identical prompt through a provider-fair adapter — identical parsing, no provider gets a structured-output crutch the other lacks.

Harness Fitscore: null
Which model works best inside Playi?

Suites that run through Playi’s own production routes, where the route may own or pin the model. This measures the system a coach actually touches, not a lab-only prompt.

Why null and not a number? Every suite’s promotion digest is still unset: no wave has passed human verification review yet. The composite score is null — not zero — for every model, on both composites. Candidate-wave tallies exist internally and are always labeled as unverified; they never feed the headline number.

Also here

Beyond the benchmark