Research

Research

Notes and audits from building the benchmark and the systems it measures. Full writeups are being prepared for publication; the underlying work is dated and summarized below so this page never claims more than exists.

In preparationwriteups pending
2026-08-14
Auditing our own multiple-choice suite for gameability

We measured the keyed answer as the strictly-longest option in 45% of an authored MCQ wave (25% chance baseline) — a classic authoring artifact — then rewrote distractors and gated null heuristics within ±10pp of chance, two-sided, so elimination strategies are caught too.

2026-08-06
Weighting a benchmark by what production actually runs

The F.I.B. weight vector is derived from Playi’s measured operation mix rather than intuition, and recorded with a date beside every published score.

The methodology itself is already public — .