Commit Graph
8 Commits
Author SHA1 Message Date
ImBenjiandClaude Opus 5 b246bd9d4b fix: a recovered replay job can no longer hijack the active run
leaseNextJob hands back any pending replay_article job, it has no idea about
runs, and the worker was attributing whatever came back to whichever run was
active. One recovered dead letter from run 1 would have been stamped with run
2's id, given run 2's feedback brief, and dragged run 2's cursor to wherever
that old article sits in the archive. A pinned run would then decide its set was
finished after a couple of articles. There are 139 dead letters and they are
built to recover, so this was not hypothetical.

The idempotency key already says which run enqueued the job. Ask it.

Also: refuse to inherit the parent's model label when starting a run. Inheriting
is exactly how run 1 came to be labelled qwen for predictions deepseek made.

The split moves to the replay container's actual restart time rather than the
commit timestamp five minutes later. Verified the running container really does
have the instrument rules, the de-anchoring and the enum before trusting it as
the boundary. It makes no difference to the partition, there are no replay
predictions at all between 15:57 and midnight that day, but the boundary should
be the thing that actually changed the prompt.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-08 01:01:36 +01:00
ImBenjiandClaude Opus 5 ec29e64e96 feat: let the generator read its own results, and measure it honestly
Nothing in the pipeline has ever fed an outcome back to the thing that makes
predictions. Calibration reads autonomy_outcomes, but calibration only gates
whether to act on a prediction, never what the prediction is. So the only thing
that has ever changed this system's output is a human editing the prompt.

score-replay-runs.js asks "compared to what". The answer is not flattering:
over 2,114 scored replay predictions the system is right 50.05% of the time
while answering "negative" to every one of the same bars scores 54.45%. It is
4.4 points below a constant, z=-4.06. The whole deficit is the prior. It says
positive on 65% of calls when 45.5% of bars beat SPY, a 19 point skew. Its
discrimination, P(up|positive) minus P(up|negative), is +3.2 points with
p=0.15, so the direction it picks is weakly informative and completely buried
by how often it defaults to positive.

The first version of that script compared each direction group's accuracy to
"always that direction" on the same rows, which is an identity and tests
nothing. T3 replaces it with the two proportion test that actually asks whether
the choice of direction carries information.

build-feedback-brief.js turns a run's scored outcomes into a memo the next run
reads before predicting. Generated from the data, not written by hand, or it is
just me editing the prompt again with extra steps.

Replay can now be pinned to an explicit article set, which is what makes two
runs comparable at all. Comparing two calendar windows of one run compares two
market regimes: the epochs in run 1 line up exactly with article vintage, E0 is
late 2024 and E2 is 2026, so nothing could be attributed. A new run also
inherits its parent's watermark instead of recomputing it from today, which
silently guaranteed a different archive slice every time.

prompt_version never moved across four material prompt changes, so every
proposal on record claims to come from the first prompt. coordinator-2 and
replay-coordinator-2.

docs/replay-run-2-preregistration.md fixes the bar before the run exists,
including which result counts as learning and which is only calibration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-08 00:56:43 +01:00
ImBenjiandClaude Opus 5 75e183aaa3 feat: tell the coordinator which instruments it can actually trade
96% of rejected proposals name something untradable: indices (SPX, DXY, ^TNX),
fx (EURUSD, XAU/USD), futures (CL=F, BZ=F) and home listings (VOW3.DE, RHM.DE,
1211.HK, 688169.SS). The analysis behind those is usually sound, it is the ticker
that cannot be used, and nothing in the prompt ever said so. We were paying for
the call and discarding the result at validation.

The rules point the model at what the allowlist actually holds: US listings and
ADRs for foreign companies, and US listed ETFs as the tradable expression of an
index, currency, rate or commodity. Every symbol named in the rules was checked
against the live allowlist first, so VWAGY, BABA, TM, SONY, SPY, QQQ, GLD, USO,
UUP and TLT all genuinely resolve. It also forbids predicting SPY itself, which
is the benchmark and whose excess return is zero by construction.

Shared between the coordinator and replay prompts rather than written twice,
since a rule that drifts between the two lanes is worse than no rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-04 19:22:48 +01:00
ImBenjiandClaude Opus 5 b8b3987e35 fix: stop the coordinator copying its own prompt example
The JSON shape in both prompts used real values as placeholders, and the model
was reading them as the answer:

  instrument: 'NVDA'          -> 397 of 611 predictions are NVDA (65%),
                                 second place is LMT with 8
  horizon_days: 10            -> 606 of 611 are horizon 10 (99.2%), out of
                                 seven allowed horizons
  direction: 'positive|negative' -> 502 of 611 are positive (82.2%)
  event_type: 'stable_enum'   -> the enum was never listed, so the model
                                 invented one label per event, 201 distinct
                                 values across 611 predictions

replayWorker had its own copy of the same prompt with the same values, which is
why both lanes show the identical skew (replay is 147/147 horizon 10, 138/147
NVDA).

Every placeholder is now a description of the field rather than a usable value,
with an explicit line saying not to copy them. event_type is validated against
the same closed family list the cohort key uses, so a label cannot mean one
thing in the prompt and another in calibration. Off-enum labels are salvaged
through the existing mapper when they are placeable and rejected when they are
not, so 'other' does not quietly become the bin again.

This does not by itself create edge. It means the next batch of predictions
measures the model's judgement instead of its willingness to copy an example.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 23:29:11 +01:00
ImBenjiandClaude Opus 5 6f1d1eee2d fix: restart the stalled autonomy pipeline and make calibration honest
Archive ingestion had been dead since 2026-08-02 because nothing in the
compose stack actually ran it. Everything downstream starved from there.

- add ingest + enrichment services. server.js only starts the scheduler when
  DURIIN_RUN_SCHEDULER is not "false", and workers/index.js was not running at
  all, so articles never got event_id/content/has_embedding and the coordinator
  had nothing to lease.
- pass an explicit origin from coordinatorWorker. it was never passed, so
  acceptProposal defaulted to 'live' and 464 historical backfill predictions
  were recorded as live. that also meant verifyEvidence got a null cutoff and
  skipped its date check entirely.
- coarsen cohortKey to event families + horizon buckets. 201 free text event
  types produced 221 cohorts averaging 2.76 samples, so the n>=30 gate could
  never be reached and everything abstained for the wrong reason.
- gate on cohort diversity, not just sample count. one ticker was roughly half
  of all resolved outcomes, so a pure count gate was measuring one company.
  unknown diversity abstains rather than passing.
- resolve the admin archive db explicitly and probe it. it relied on a
  Dockerfile symlink, and without it better-sqlite3 quietly creates an empty
  file and serves a phantom archive.
- clamp implausible future publication dates at ingest.
- pin the db backend to sqlite by default. compose hardcoded postgres "true",
  which would have overridden the operator's own .env on the next redeploy and
  pointed everything at a stale snapshot.

scripts/repair-autonomy-labels.js relabels the affected rows. it is dry run by
default and has not been applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 21:43:24 +01:00
ImBenji f0b598a3b8 feat: support postgres autonomy runtime 2026-08-17 13:43:50 +01:00
ImBenji 2c023c8962 fix: let replay recover from dead letter jobs 2026-08-08 23:32:41 +01:00
ImBenji 5877783862 feat: add isolated historical replay calibration 2026-08-04 22:00:11 +01:00