Research
Research
Notes and audits from building the benchmark and the systems it measures. Full writeups are being prepared for publication; the underlying work is dated and summarized below so this page never claims more than exists.
In preparationwriteups pending
2026-08-14
Auditing our own multiple-choice suite for gameability
We measured the keyed answer as the strictly-longest option in 45% of an authored MCQ wave (25% chance baseline) — a classic authoring artifact — then rewrote distractors and gated null heuristics within ±10pp of chance, two-sided, so elimination strategies are caught too.
2026-08-06
Weighting a benchmark by what production actually runs
The F.I.B. weight vector is derived from Playi’s measured operation mix rather than intuition, and recorded with a date beside every published score.
The methodology itself is already public — .