diff --git a/docs/replay-run-2-preregistration.md b/docs/replay-run-2-preregistration.md index 5ff365c..93d4dbd 100644 --- a/docs/replay-run-2-preregistration.md +++ b/docs/replay-run-2-preregistration.md @@ -88,6 +88,21 @@ model the base rate will pull it toward the base rate. So: Beating run 1 while still sitting below 57.39% is not a system worth trading. That distinction gets reported every time, not just when it is convenient. +## Yield is an outcome too, declared before it is known + +The brief tells the model that an empty predictions array is always available +and that it should prefer one on the families it reads worst. So run 2 may +answer fewer than the 716 articles run 1 answered. Run 1's yield on this set is +100% by construction, since these are precisely the articles it answered. + +Recorded now, with two run-2 proposals on the board and no idea what the rate +will be: increased selectivity is a real behavioural change and gets reported as +one, not quietly dropped for shrinking the sample. T4 runs on the shared +articles whatever that number turns out to be. If the shared set falls below +about 300 articles the paired test loses the power to see a 5 point shift, and +the honest report is then "the feedback made it far more selective and the +sample it left is too small to grade", not a null result dressed up as one. + ## Housekeeping that is easy to forget Run 2 stays `running` once it exhausts its 716 articles, and the replay lane