Files
Duriin-API/docs/replay-run-2-preregistration.md
T
ImBenjiandClaude Opus 5 b246bd9d4b fix: a recovered replay job can no longer hijack the active run
leaseNextJob hands back any pending replay_article job, it has no idea about
runs, and the worker was attributing whatever came back to whichever run was
active. One recovered dead letter from run 1 would have been stamped with run
2's id, given run 2's feedback brief, and dragged run 2's cursor to wherever
that old article sits in the archive. A pinned run would then decide its set was
finished after a couple of articles. There are 139 dead letters and they are
built to recover, so this was not hypothetical.

The idempotency key already says which run enqueued the job. Ask it.

Also: refuse to inherit the parent's model label when starting a run. Inheriting
is exactly how run 1 came to be labelled qwen for predictions deepseek made.

The split moves to the replay container's actual restart time rather than the
commit timestamp five minutes later. Verified the running container really does
have the instrument rules, the de-anchoring and the enum before trusting it as
the boundary. It makes no difference to the partition, there are no replay
predictions at all between 15:57 and midnight that day, but the boundary should
be the thing that actually changed the prompt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-08 01:01:36 +01:00

5.0 KiB

Replay run 2: pre-registration

Written 2026-09-08, before run 2 exists. The numbers below are run 1's, measured on the evaluation slice only. They are fixed. If the analysis after run 2 uses a different bar than the one written here, the analysis is wrong, not the bar.

What is being tested

Whether feeding the generator its own scored results changes what it predicts, and whether the change is an improvement.

Nothing in the pipeline has ever read autonomy_outcomes back into the thing that makes predictions. Calibration reads outcomes, but calibration only decides whether to ACT on a prediction, never what the prediction is. So the only thing that has ever altered this system's output is a human editing the prompt. Run 2 is the first time the system is told about its own mistakes.

Design

One variable. Run 2 answers exactly the articles run 1 answered under the current prompt and the current model, with the same model, same prompt, plus a feedback brief generated by scripts/build-feedback-brief.js.

  • Split on created_at >= "2026-09-04 19:17:43", when duriin-api-replay-1 restarted onto the prompt it still runs today. That is the container start time, not the commit timestamp, which is five minutes later and would have put a handful of old-prompt predictions on the new-prompt side. Verified directly: the running container has the instrument rules, the de-anchored placeholders and the event_type enum in /app/workers/replayWorker.js.
  • There is no contamination to argue about. Replay was between daily budgets across the restart, so no replay prediction exists between 15:57 on 09-04 and 00:00 on 09-05. Splits at 19:17:43, at 19:30 and at midnight all produce the identical partition, 1,142 training and 984 evaluation.
  • The brief is derived from the 1,133 training predictions only. No evaluation row contributes a single number to the text. Deriving the lesson and grading it on the same rows would measure nothing.
  • Evaluation set: 716 articles, 984 run-1 predictions, all ~deepseek/deepseek-v4-flash-latest. 713 of those articles have a scored run-1 prediction and are the pairable set; the other three are replayed but cannot enter T4.
  • The 1,133 scored training predictions span two models, roughly 627 qwen and 506 deepseek. So the brief describes the mistakes of the system as it has been, not of deepseek alone. Run 2 is deepseek throughout, as is the run-1 half it is measured against.

Run 1 on the evaluation slice, the bar

metric run 1
predictions scored 981 over 713 articles
accuracy 50.56%
always_negative on the same bars 57.39%
edge over the constant -6.83 points
signed excess, system 0.679%
signed excess, always_negative 1.264%
discrimination P(up given positive) 44.35% (n=593)
discrimination P(up given negative) 39.95% (n=388)
discrimination spread +4.40 points, z=1.363, p=0.173
share of calls that were positive 60.45%, against 42.61% of bars up

Tests, declared now

  • Primary, T4. Paired per-article accuracy, run 2 minus run 1, over the shared articles. Two sided paired t. Per article, not per prediction, so one article that produced eleven calls does not outvote one that produced a single call.
  • Secondary, T3. Discrimination spread. Run 1 is +4.40 points.
  • Absolute, T1. Run 2 accuracy against always_negative, 57.39%.

What each outcome means, declared now

The brief tells the model its positive share is 17.8 points too high. Telling a model the base rate will pull it toward the base rate. So:

  • Accuracy up, discrimination spread flat. The expected result. This is calibration, not skill. The system learned its prior was wrong, which is worth having and is genuinely the loop working, but it is not evidence that it reads news any better. Do not report it as new skill.
  • Accuracy up AND discrimination spread up, T3 significant. Genuine learning. The feedback changed which way it calls things, not just how often. This is the only result that justifies building the loop into the workers.
  • T4 flat. The feedback changed nothing. Either the brief is too weak to move the model or the model cannot use this kind of instruction. Either way the answer to "should the loop be automated" is no, and the structural alternatives become the next move.
  • T4 negative. The feedback made it worse. Report it as such and stop.

Beating run 1 while still sitting below 57.39% is not a system worth trading. That distinction gets reported every time, not just when it is convenient.

Housekeeping that is easy to forget

Run 2 stays running once it exhausts its 716 articles, and the replay lane just idles. That is intended, it keeps the spend at zero while the outcomes mature. Mark it complete when the results are read, otherwise it becomes the same stale metadata run 1 carried for a month. But do not mark it complete before reading, because activeRun would immediately create run 3 with no pinned set and no brief and start walking all 23k articles again.