feat: let the generator read its own results, and measure it honestly

Nothing in the pipeline has ever fed an outcome back to the thing that makes
predictions. Calibration reads autonomy_outcomes, but calibration only gates
whether to act on a prediction, never what the prediction is. So the only thing
that has ever changed this system's output is a human editing the prompt.

score-replay-runs.js asks "compared to what". The answer is not flattering:
over 2,114 scored replay predictions the system is right 50.05% of the time
while answering "negative" to every one of the same bars scores 54.45%. It is
4.4 points below a constant, z=-4.06. The whole deficit is the prior. It says
positive on 65% of calls when 45.5% of bars beat SPY, a 19 point skew. Its
discrimination, P(up|positive) minus P(up|negative), is +3.2 points with
p=0.15, so the direction it picks is weakly informative and completely buried
by how often it defaults to positive.

The first version of that script compared each direction group's accuracy to
"always that direction" on the same rows, which is an identity and tests
nothing. T3 replaces it with the two proportion test that actually asks whether
the choice of direction carries information.

build-feedback-brief.js turns a run's scored outcomes into a memo the next run
reads before predicting. Generated from the data, not written by hand, or it is
just me editing the prompt again with extra steps.

Replay can now be pinned to an explicit article set, which is what makes two
runs comparable at all. Comparing two calendar windows of one run compares two
market regimes: the epochs in run 1 line up exactly with article vintage, E0 is
late 2024 and E2 is 2026, so nothing could be attributed. A new run also
inherits its parent's watermark instead of recomputing it from today, which
silently guaranteed a different archive slice every time.

prompt_version never moved across four material prompt changes, so every
proposal on record claims to come from the first prompt. coordinator-2 and
replay-coordinator-2.

docs/replay-run-2-preregistration.md fixes the bar before the run exists,
including which result counts as learning and which is only calibration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
This commit is contained in:
ImBenji
2026-09-08 00:56:43 +01:00
co-authored by Claude Opus 5
parent 028c07688d
commit ec29e64e96
8 changed files with 881 additions and 12 deletions
+17
View File
@@ -209,6 +209,19 @@ function initAutonomySchema(db) {
CREATE INDEX IF NOT EXISTS idx_autonomy_replay_runs_active
ON autonomy_replay_runs(status, cursor_article_id);
-- An explicit article set for a run. When a run has rows here the scheduler
-- walks exactly these and nothing else, which is the only way to point two
-- runs at the same evidence. Runs without rows here keep walking the
-- archive by cursor exactly as before.
CREATE TABLE IF NOT EXISTS autonomy_replay_run_articles (
run_id INTEGER NOT NULL REFERENCES autonomy_replay_runs(id),
article_id INTEGER NOT NULL,
effective_at TEXT,
PRIMARY KEY (run_id, article_id)
);
CREATE INDEX IF NOT EXISTS idx_autonomy_replay_run_articles_walk
ON autonomy_replay_run_articles(run_id, effective_at, article_id);
CREATE TABLE IF NOT EXISTS autonomy_replay_evaluations (
prediction_id INTEGER PRIMARY KEY REFERENCES autonomy_predictions(id),
replay_run_id INTEGER NOT NULL REFERENCES autonomy_replay_runs(id),
@@ -240,6 +253,10 @@ function initAutonomySchema(db) {
"ALTER TABLE autonomy_predictions ADD COLUMN origin TEXT NOT NULL DEFAULT 'live'",
'ALTER TABLE autonomy_predictions ADD COLUMN replay_run_id INTEGER',
'ALTER TABLE autonomy_replay_runs ADD COLUMN cursor_effective_at TEXT',
// what this run was told about its predecessor's mistakes, kept on the run
// so a result can always be traced back to the text that produced it
'ALTER TABLE autonomy_replay_runs ADD COLUMN feedback_brief TEXT',
'ALTER TABLE autonomy_replay_runs ADD COLUMN parent_run_id INTEGER',
"ALTER TABLE autonomy_calibration_snapshots ADD COLUMN source TEXT NOT NULL DEFAULT 'live'",
'ALTER TABLE autonomy_calibration_snapshots ADD COLUMN replay_run_id INTEGER',
'ALTER TABLE autonomy_calibration_snapshots ADD COLUMN distinct_instruments INTEGER',