How We Benchmarked Recorded-Video Verification, and What the Result Does Not Say
38 clips across 14 titles and no wrong automatic settles. Here is how the offline benchmark was run, and where its limits are.
Phase 1 of the Outcomes Oracle verifies game results from recorded video. Before saying anything about it, we ran an offline benchmark and published the result on the status page. This post explains the method, because the number on its own is easy to misread.
Each clip was run through the full Vision pipeline from scratch, with no fallback. Each clip has a ground-truth outcome that a person checked. A hold counts as correct only where a hold was the expected result, so a pipeline that simply holds everything does not score well.
The headline check is the wrong auto-settle: an automatic settle with an incorrect outcome. It is the failure that matters most, because it moves an event to settled on bad evidence. Across 38 clips in 14 titles, all 38 matched the expected result and there were no wrong auto-settles. Of those clips, 31 settled automatically, 4 were held as expected, 1 was sent for confirmation and 2 were voided.
The benchmark carries its own label because of its limits. The clips are public and were chosen by the team, 2 to 5 per title, so this is not a random sample. It is recorded video only; there is no live-video result in the set. It shows the pipeline behaving on these clips. It is not an accuracy rate for any title and not a production test.
That is why the coverage page says benchmarked rather than tested. Tested is reserved for authenticated production tests, and none are on record yet. They will appear on the status page when a named partner export lands.
The figures above are as of 2026-08-10.