Files
Duriin-API/docs/replay-run-2-preregistration.md
T
ImBenjiandClaude Opus 5 20fb021983 docs: declare run 2 answer rate as an outcome before it is known
The brief tells the model an empty predictions array is always available, so
run 2 may answer fewer articles than run 1 did. Writing down now, while there
are two run-2 proposals on the board, that selectivity gets reported as a
result rather than quietly treated as a smaller sample.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-08 01:55:38 +01:00

114 lines
5.9 KiB
Markdown

# Replay run 2: pre-registration
Written 2026-09-08, before run 2 exists. The numbers below are run 1's, measured
on the evaluation slice only. They are fixed. If the analysis after run 2 uses a
different bar than the one written here, the analysis is wrong, not the bar.
## What is being tested
Whether feeding the generator its own scored results changes what it predicts,
and whether the change is an improvement.
Nothing in the pipeline has ever read `autonomy_outcomes` back into the thing
that makes predictions. Calibration reads outcomes, but calibration only decides
whether to ACT on a prediction, never what the prediction is. So the only thing
that has ever altered this system's output is a human editing the prompt. Run 2
is the first time the system is told about its own mistakes.
## Design
One variable. Run 2 answers exactly the articles run 1 answered under the
current prompt and the current model, with the same model, same prompt, plus a
feedback brief generated by `scripts/build-feedback-brief.js`.
- Split on `created_at >= "2026-09-04 19:17:43"`, when `duriin-api-replay-1`
restarted onto the prompt it still runs today. That is the container start
time, not the commit timestamp, which is five minutes later and would have put
a handful of old-prompt predictions on the new-prompt side. Verified directly:
the running container has the instrument rules, the de-anchored placeholders
and the event_type enum in `/app/workers/replayWorker.js`.
- There is no contamination to argue about. Replay was between daily budgets
across the restart, so no replay prediction exists between 15:57 on 09-04 and
00:00 on 09-05. Splits at 19:17:43, at 19:30 and at midnight all produce the
identical partition, 1,142 training and 984 evaluation.
- The brief is derived from the 1,133 training predictions only. No evaluation
row contributes a single number to the text. Deriving the lesson and grading it
on the same rows would measure nothing.
- Evaluation set: 716 articles, 984 run-1 predictions, all
`~deepseek/deepseek-v4-flash-latest`. 713 of those articles have a scored
run-1 prediction and are the pairable set; the other three are replayed but
cannot enter T4.
- The 1,133 scored training predictions span two models, roughly 627 qwen and
506 deepseek. So the brief describes the mistakes of the system as it has
been, not of deepseek alone. Run 2 is deepseek throughout, as is the run-1
half it is measured against.
## Run 1 on the evaluation slice, the bar
| metric | run 1 |
| --- | --- |
| predictions scored | 981 over 713 articles |
| accuracy | 50.56% |
| always_negative on the same bars | 57.39% |
| edge over the constant | **-6.83 points** |
| signed excess, system | 0.679% |
| signed excess, always_negative | 1.264% |
| discrimination P(up given positive) | 44.35% (n=593) |
| discrimination P(up given negative) | 39.95% (n=388) |
| discrimination spread | **+4.40 points**, z=1.363, p=0.173 |
| share of calls that were positive | 60.45%, against 42.61% of bars up |
## Tests, declared now
- **Primary, T4.** Paired per-article accuracy, run 2 minus run 1, over the
shared articles. Two sided paired t. Per article, not per prediction, so one
article that produced eleven calls does not outvote one that produced a single
call.
- **Secondary, T3.** Discrimination spread. Run 1 is +4.40 points.
- **Absolute, T1.** Run 2 accuracy against always_negative, 57.39%.
## What each outcome means, declared now
The brief tells the model its positive share is 17.8 points too high. Telling a
model the base rate will pull it toward the base rate. So:
- **Accuracy up, discrimination spread flat.** The expected result. This is
calibration, not skill. The system learned its prior was wrong, which is worth
having and is genuinely the loop working, but it is not evidence that it reads
news any better. Do not report it as new skill.
- **Accuracy up AND discrimination spread up, T3 significant.** Genuine
learning. The feedback changed which way it calls things, not just how often.
This is the only result that justifies building the loop into the workers.
- **T4 flat.** The feedback changed nothing. Either the brief is too weak to move
the model or the model cannot use this kind of instruction. Either way the
answer to "should the loop be automated" is no, and the structural
alternatives become the next move.
- **T4 negative.** The feedback made it worse. Report it as such and stop.
Beating run 1 while still sitting below 57.39% is not a system worth trading.
That distinction gets reported every time, not just when it is convenient.
## Yield is an outcome too, declared before it is known
The brief tells the model that an empty predictions array is always available
and that it should prefer one on the families it reads worst. So run 2 may
answer fewer than the 716 articles run 1 answered. Run 1's yield on this set is
100% by construction, since these are precisely the articles it answered.
Recorded now, with two run-2 proposals on the board and no idea what the rate
will be: increased selectivity is a real behavioural change and gets reported as
one, not quietly dropped for shrinking the sample. T4 runs on the shared
articles whatever that number turns out to be. If the shared set falls below
about 300 articles the paired test loses the power to see a 5 point shift, and
the honest report is then "the feedback made it far more selective and the
sample it left is too small to grade", not a null result dressed up as one.
## Housekeeping that is easy to forget
Run 2 stays `running` once it exhausts its 716 articles, and the replay lane
just idles. That is intended, it keeps the spend at zero while the outcomes
mature. Mark it `complete` when the results are read, otherwise it becomes the
same stale metadata run 1 carried for a month. But do not mark it complete
before reading, because `activeRun` would immediately create run 3 with no
pinned set and no brief and start walking all 23k articles again.