Because the proposal and human decision are recorded before the outcome is known, and the outcome is appended later as a separate record, the three can be compared afterwards. The counting is deliberately the least clever part of the system, and it refuses more often than it reports.
min_sample = 20 resolved outcomes
“This analyst has overridden LyraMind 17 times. 11 resolved in their direction and 6 in LyraMind's.”
LyraMind does not yet say this analyst is more accurate than the system, or that one assumption category performs better than another. The sample is too small.
An example of the form of sentence the record supports, and of the one it declines to produce. The names and counts are invented; no customer record is described here.
That sentence is only available because something recorded the pre-outcome state before the outcome was known. It cannot be reconstructed from an archive, and it cannot be backfilled, which is why the record is built this way from the first decision rather than added later.
A comparison between two groups carries its own, higher floor and a robustness test on top of it. “Regulatory overrides perform better than demand overrides” is a claim about two samples at once, so it is reported only when both clear the floor and the gap survives the test.