The brief tells the model an empty predictions array is always available, so run 2 may answer fewer articles than run 1 did. Writing down now, while there are two run-2 proposals on the board, that selectivity gets reported as a result rather than quietly treated as a smaller sample. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
114 lines
5.9 KiB
Markdown
114 lines
5.9 KiB
Markdown
# Replay run 2: pre-registration
|
|
|
|
Written 2026-09-08, before run 2 exists. The numbers below are run 1's, measured
|
|
on the evaluation slice only. They are fixed. If the analysis after run 2 uses a
|
|
different bar than the one written here, the analysis is wrong, not the bar.
|
|
|
|
## What is being tested
|
|
|
|
Whether feeding the generator its own scored results changes what it predicts,
|
|
and whether the change is an improvement.
|
|
|
|
Nothing in the pipeline has ever read `autonomy_outcomes` back into the thing
|
|
that makes predictions. Calibration reads outcomes, but calibration only decides
|
|
whether to ACT on a prediction, never what the prediction is. So the only thing
|
|
that has ever altered this system's output is a human editing the prompt. Run 2
|
|
is the first time the system is told about its own mistakes.
|
|
|
|
## Design
|
|
|
|
One variable. Run 2 answers exactly the articles run 1 answered under the
|
|
current prompt and the current model, with the same model, same prompt, plus a
|
|
feedback brief generated by `scripts/build-feedback-brief.js`.
|
|
|
|
- Split on `created_at >= "2026-09-04 19:17:43"`, when `duriin-api-replay-1`
|
|
restarted onto the prompt it still runs today. That is the container start
|
|
time, not the commit timestamp, which is five minutes later and would have put
|
|
a handful of old-prompt predictions on the new-prompt side. Verified directly:
|
|
the running container has the instrument rules, the de-anchored placeholders
|
|
and the event_type enum in `/app/workers/replayWorker.js`.
|
|
- There is no contamination to argue about. Replay was between daily budgets
|
|
across the restart, so no replay prediction exists between 15:57 on 09-04 and
|
|
00:00 on 09-05. Splits at 19:17:43, at 19:30 and at midnight all produce the
|
|
identical partition, 1,142 training and 984 evaluation.
|
|
- The brief is derived from the 1,133 training predictions only. No evaluation
|
|
row contributes a single number to the text. Deriving the lesson and grading it
|
|
on the same rows would measure nothing.
|
|
- Evaluation set: 716 articles, 984 run-1 predictions, all
|
|
`~deepseek/deepseek-v4-flash-latest`. 713 of those articles have a scored
|
|
run-1 prediction and are the pairable set; the other three are replayed but
|
|
cannot enter T4.
|
|
- The 1,133 scored training predictions span two models, roughly 627 qwen and
|
|
506 deepseek. So the brief describes the mistakes of the system as it has
|
|
been, not of deepseek alone. Run 2 is deepseek throughout, as is the run-1
|
|
half it is measured against.
|
|
|
|
## Run 1 on the evaluation slice, the bar
|
|
|
|
| metric | run 1 |
|
|
| --- | --- |
|
|
| predictions scored | 981 over 713 articles |
|
|
| accuracy | 50.56% |
|
|
| always_negative on the same bars | 57.39% |
|
|
| edge over the constant | **-6.83 points** |
|
|
| signed excess, system | 0.679% |
|
|
| signed excess, always_negative | 1.264% |
|
|
| discrimination P(up given positive) | 44.35% (n=593) |
|
|
| discrimination P(up given negative) | 39.95% (n=388) |
|
|
| discrimination spread | **+4.40 points**, z=1.363, p=0.173 |
|
|
| share of calls that were positive | 60.45%, against 42.61% of bars up |
|
|
|
|
## Tests, declared now
|
|
|
|
- **Primary, T4.** Paired per-article accuracy, run 2 minus run 1, over the
|
|
shared articles. Two sided paired t. Per article, not per prediction, so one
|
|
article that produced eleven calls does not outvote one that produced a single
|
|
call.
|
|
- **Secondary, T3.** Discrimination spread. Run 1 is +4.40 points.
|
|
- **Absolute, T1.** Run 2 accuracy against always_negative, 57.39%.
|
|
|
|
## What each outcome means, declared now
|
|
|
|
The brief tells the model its positive share is 17.8 points too high. Telling a
|
|
model the base rate will pull it toward the base rate. So:
|
|
|
|
- **Accuracy up, discrimination spread flat.** The expected result. This is
|
|
calibration, not skill. The system learned its prior was wrong, which is worth
|
|
having and is genuinely the loop working, but it is not evidence that it reads
|
|
news any better. Do not report it as new skill.
|
|
- **Accuracy up AND discrimination spread up, T3 significant.** Genuine
|
|
learning. The feedback changed which way it calls things, not just how often.
|
|
This is the only result that justifies building the loop into the workers.
|
|
- **T4 flat.** The feedback changed nothing. Either the brief is too weak to move
|
|
the model or the model cannot use this kind of instruction. Either way the
|
|
answer to "should the loop be automated" is no, and the structural
|
|
alternatives become the next move.
|
|
- **T4 negative.** The feedback made it worse. Report it as such and stop.
|
|
|
|
Beating run 1 while still sitting below 57.39% is not a system worth trading.
|
|
That distinction gets reported every time, not just when it is convenient.
|
|
|
|
## Yield is an outcome too, declared before it is known
|
|
|
|
The brief tells the model that an empty predictions array is always available
|
|
and that it should prefer one on the families it reads worst. So run 2 may
|
|
answer fewer than the 716 articles run 1 answered. Run 1's yield on this set is
|
|
100% by construction, since these are precisely the articles it answered.
|
|
|
|
Recorded now, with two run-2 proposals on the board and no idea what the rate
|
|
will be: increased selectivity is a real behavioural change and gets reported as
|
|
one, not quietly dropped for shrinking the sample. T4 runs on the shared
|
|
articles whatever that number turns out to be. If the shared set falls below
|
|
about 300 articles the paired test loses the power to see a 5 point shift, and
|
|
the honest report is then "the feedback made it far more selective and the
|
|
sample it left is too small to grade", not a null result dressed up as one.
|
|
|
|
## Housekeeping that is easy to forget
|
|
|
|
Run 2 stays `running` once it exhausts its 716 articles, and the replay lane
|
|
just idles. That is intended, it keeps the spend at zero while the outcomes
|
|
mature. Mark it `complete` when the results are read, otherwise it becomes the
|
|
same stale metadata run 1 carried for a month. But do not mark it complete
|
|
before reading, because `activeRun` would immediately create run 3 with no
|
|
pinned set and no brief and start walking all 23k articles again.
|