Nothing in the pipeline has ever fed an outcome back to the thing that makes predictions. Calibration reads autonomy_outcomes, but calibration only gates whether to act on a prediction, never what the prediction is. So the only thing that has ever changed this system's output is a human editing the prompt. score-replay-runs.js asks "compared to what". The answer is not flattering: over 2,114 scored replay predictions the system is right 50.05% of the time while answering "negative" to every one of the same bars scores 54.45%. It is 4.4 points below a constant, z=-4.06. The whole deficit is the prior. It says positive on 65% of calls when 45.5% of bars beat SPY, a 19 point skew. Its discrimination, P(up|positive) minus P(up|negative), is +3.2 points with p=0.15, so the direction it picks is weakly informative and completely buried by how often it defaults to positive. The first version of that script compared each direction group's accuracy to "always that direction" on the same rows, which is an identity and tests nothing. T3 replaces it with the two proportion test that actually asks whether the choice of direction carries information. build-feedback-brief.js turns a run's scored outcomes into a memo the next run reads before predicting. Generated from the data, not written by hand, or it is just me editing the prompt again with extra steps. Replay can now be pinned to an explicit article set, which is what makes two runs comparable at all. Comparing two calendar windows of one run compares two market regimes: the epochs in run 1 line up exactly with article vintage, E0 is late 2024 and E2 is 2026, so nothing could be attributed. A new run also inherits its parent's watermark instead of recomputing it from today, which silently guaranteed a different archive slice every time. prompt_version never moved across four material prompt changes, so every proposal on record claims to come from the first prompt. coordinator-2 and replay-coordinator-2. docs/replay-run-2-preregistration.md fixes the bar before the run exists, including which result counts as learning and which is only calibration. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
77 lines
3.6 KiB
Markdown
77 lines
3.6 KiB
Markdown
# Replay run 2: pre-registration
|
|
|
|
Written 2026-09-08, before run 2 exists. The numbers below are run 1's, measured
|
|
on the evaluation slice only. They are fixed. If the analysis after run 2 uses a
|
|
different bar than the one written here, the analysis is wrong, not the bar.
|
|
|
|
## What is being tested
|
|
|
|
Whether feeding the generator its own scored results changes what it predicts,
|
|
and whether the change is an improvement.
|
|
|
|
Nothing in the pipeline has ever read `autonomy_outcomes` back into the thing
|
|
that makes predictions. Calibration reads outcomes, but calibration only decides
|
|
whether to ACT on a prediction, never what the prediction is. So the only thing
|
|
that has ever altered this system's output is a human editing the prompt. Run 2
|
|
is the first time the system is told about its own mistakes.
|
|
|
|
## Design
|
|
|
|
One variable. Run 2 answers exactly the articles run 1 answered under the
|
|
current prompt and the current model, with the same model, same prompt, plus a
|
|
feedback brief generated by `scripts/build-feedback-brief.js`.
|
|
|
|
- Split on `created_at >= "2026-09-04 19:30"`, the deploy that put the instrument
|
|
rules into the replay prompt. Everything before it is TRAINING, everything at
|
|
or after it is EVALUATION.
|
|
- The brief is derived from the 1,133 training predictions only. No evaluation
|
|
row contributes a single number to the text. Deriving the lesson and grading it
|
|
on the same rows would measure nothing.
|
|
- Evaluation set: 713 articles, 981 run-1 predictions, all
|
|
`~deepseek/deepseek-v4-flash-latest`.
|
|
|
|
## Run 1 on the evaluation slice, the bar
|
|
|
|
| metric | run 1 |
|
|
| --- | --- |
|
|
| predictions scored | 981 over 713 articles |
|
|
| accuracy | 50.56% |
|
|
| always_negative on the same bars | 57.39% |
|
|
| edge over the constant | **-6.83 points** |
|
|
| signed excess, system | 0.679% |
|
|
| signed excess, always_negative | 1.264% |
|
|
| discrimination P(up given positive) | 44.35% (n=593) |
|
|
| discrimination P(up given negative) | 39.95% (n=388) |
|
|
| discrimination spread | **+4.40 points**, z=1.363, p=0.173 |
|
|
| share of calls that were positive | 60.45%, against 42.61% of bars up |
|
|
|
|
## Tests, declared now
|
|
|
|
- **Primary, T4.** Paired per-article accuracy, run 2 minus run 1, over the
|
|
shared articles. Two sided paired t. Per article, not per prediction, so one
|
|
article that produced eleven calls does not outvote one that produced a single
|
|
call.
|
|
- **Secondary, T3.** Discrimination spread. Run 1 is +4.40 points.
|
|
- **Absolute, T1.** Run 2 accuracy against always_negative, 57.39%.
|
|
|
|
## What each outcome means, declared now
|
|
|
|
The brief tells the model its positive share is 17.8 points too high. Telling a
|
|
model the base rate will pull it toward the base rate. So:
|
|
|
|
- **Accuracy up, discrimination spread flat.** The expected result. This is
|
|
calibration, not skill. The system learned its prior was wrong, which is worth
|
|
having and is genuinely the loop working, but it is not evidence that it reads
|
|
news any better. Do not report it as new skill.
|
|
- **Accuracy up AND discrimination spread up, T3 significant.** Genuine
|
|
learning. The feedback changed which way it calls things, not just how often.
|
|
This is the only result that justifies building the loop into the workers.
|
|
- **T4 flat.** The feedback changed nothing. Either the brief is too weak to move
|
|
the model or the model cannot use this kind of instruction. Either way the
|
|
answer to "should the loop be automated" is no, and the structural
|
|
alternatives become the next move.
|
|
- **T4 negative.** The feedback made it worse. Report it as such and stop.
|
|
|
|
Beating run 1 while still sitting below 57.39% is not a system worth trading.
|
|
That distinction gets reported every time, not just when it is convenient.
|