Nothing in the pipeline has ever fed an outcome back to the thing that makes
predictions. Calibration reads autonomy_outcomes, but calibration only gates
whether to act on a prediction, never what the prediction is. So the only thing
that has ever changed this system's output is a human editing the prompt.
score-replay-runs.js asks "compared to what". The answer is not flattering:
over 2,114 scored replay predictions the system is right 50.05% of the time
while answering "negative" to every one of the same bars scores 54.45%. It is
4.4 points below a constant, z=-4.06. The whole deficit is the prior. It says
positive on 65% of calls when 45.5% of bars beat SPY, a 19 point skew. Its
discrimination, P(up|positive) minus P(up|negative), is +3.2 points with
p=0.15, so the direction it picks is weakly informative and completely buried
by how often it defaults to positive.
The first version of that script compared each direction group's accuracy to
"always that direction" on the same rows, which is an identity and tests
nothing. T3 replaces it with the two proportion test that actually asks whether
the choice of direction carries information.
build-feedback-brief.js turns a run's scored outcomes into a memo the next run
reads before predicting. Generated from the data, not written by hand, or it is
just me editing the prompt again with extra steps.
Replay can now be pinned to an explicit article set, which is what makes two
runs comparable at all. Comparing two calendar windows of one run compares two
market regimes: the epochs in run 1 line up exactly with article vintage, E0 is
late 2024 and E2 is 2026, so nothing could be attributed. A new run also
inherits its parent's watermark instead of recomputing it from today, which
silently guaranteed a different archive slice every time.
prompt_version never moved across four material prompt changes, so every
proposal on record claims to come from the first prompt. coordinator-2 and
replay-coordinator-2.
docs/replay-run-2-preregistration.md fixes the bar before the run exists,
including which result counts as learning and which is only calibration.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The graph has been built for months and fed nothing but a dashboard. Nothing in
the autonomy pipeline has ever read an edge: not the coordinator, not
calibration, not execution. Its only downstream consumer, trade_signals, last
produced anything in April. It was roughly 70% of the llm bill and informed no
prediction, decision or order.
Relationships are the one piece of context a per-event coordinator genuinely
cannot derive from its own articles, because "this company supplies that one" is
knowledge about companies rather than about this event. So the coordinator now
receives the relationships of the companies the event is about, and is told to
name the relationship in causal_channel when it reasons through one.
The cutoff filter is the part that matters. first_seen_at on a relationship is
derived from article dates rather than processing time, so a historical proposal
only sees what the world had actually revealed by its own cutoff. Without that
this feature would quietly reintroduce the lookahead the evidence check exists to
prevent, and it is pinned by a test rather than left to review.
Relationships are explicitly background rather than evidence: predictions still
have to cite the article ids the story came from, and an instrument the articles
give no reason to care about is still not a prediction.
strategy_version moves to autonomy-2, because a prompt change this material
changes what a prediction means and the two populations should be comparable
later rather than silently blended.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
96% of rejected proposals name something untradable: indices (SPX, DXY, ^TNX),
fx (EURUSD, XAU/USD), futures (CL=F, BZ=F) and home listings (VOW3.DE, RHM.DE,
1211.HK, 688169.SS). The analysis behind those is usually sound, it is the ticker
that cannot be used, and nothing in the prompt ever said so. We were paying for
the call and discarding the result at validation.
The rules point the model at what the allowlist actually holds: US listings and
ADRs for foreign companies, and US listed ETFs as the tradable expression of an
index, currency, rate or commodity. Every symbol named in the rules was checked
against the live allowlist first, so VWAGY, BABA, TM, SONY, SPY, QQQ, GLD, USO,
UUP and TLT all genuinely resolve. It also forbids predicting SPY itself, which
is the benchmark and whose excess return is zero by construction.
Shared between the coordinator and replay prompts rather than written twice,
since a rule that drifts between the two lanes is worse than no rule.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The JSON shape in both prompts used real values as placeholders, and the model
was reading them as the answer:
instrument: 'NVDA' -> 397 of 611 predictions are NVDA (65%),
second place is LMT with 8
horizon_days: 10 -> 606 of 611 are horizon 10 (99.2%), out of
seven allowed horizons
direction: 'positive|negative' -> 502 of 611 are positive (82.2%)
event_type: 'stable_enum' -> the enum was never listed, so the model
invented one label per event, 201 distinct
values across 611 predictions
replayWorker had its own copy of the same prompt with the same values, which is
why both lanes show the identical skew (replay is 147/147 horizon 10, 138/147
NVDA).
Every placeholder is now a description of the field rather than a usable value,
with an explicit line saying not to copy them. event_type is validated against
the same closed family list the cohort key uses, so a label cannot mean one
thing in the prompt and another in calibration. Off-enum labels are salvaged
through the existing mapper when they are placeable and rejected when they are
not, so 'other' does not quietly become the bin again.
This does not by itself create edge. It means the next batch of predictions
measures the model's judgement instead of its willingness to copy an example.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Archive ingestion had been dead since 2026-08-02 because nothing in the
compose stack actually ran it. Everything downstream starved from there.
- add ingest + enrichment services. server.js only starts the scheduler when
DURIIN_RUN_SCHEDULER is not "false", and workers/index.js was not running at
all, so articles never got event_id/content/has_embedding and the coordinator
had nothing to lease.
- pass an explicit origin from coordinatorWorker. it was never passed, so
acceptProposal defaulted to 'live' and 464 historical backfill predictions
were recorded as live. that also meant verifyEvidence got a null cutoff and
skipped its date check entirely.
- coarsen cohortKey to event families + horizon buckets. 201 free text event
types produced 221 cohorts averaging 2.76 samples, so the n>=30 gate could
never be reached and everything abstained for the wrong reason.
- gate on cohort diversity, not just sample count. one ticker was roughly half
of all resolved outcomes, so a pure count gate was measuring one company.
unknown diversity abstains rather than passing.
- resolve the admin archive db explicitly and probe it. it relied on a
Dockerfile symlink, and without it better-sqlite3 quietly creates an empty
file and serves a phantom archive.
- clamp implausible future publication dates at ingest.
- pin the db backend to sqlite by default. compose hardcoded postgres "true",
which would have overridden the operator's own .env on the next redeploy and
pointed everything at a stale snapshot.
scripts/repair-autonomy-labels.js relabels the affected rows. it is dry run by
default and has not been applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb