The graph has been built for months and fed nothing but a dashboard. Nothing in
the autonomy pipeline has ever read an edge: not the coordinator, not
calibration, not execution. Its only downstream consumer, trade_signals, last
produced anything in April. It was roughly 70% of the llm bill and informed no
prediction, decision or order.
Relationships are the one piece of context a per-event coordinator genuinely
cannot derive from its own articles, because "this company supplies that one" is
knowledge about companies rather than about this event. So the coordinator now
receives the relationships of the companies the event is about, and is told to
name the relationship in causal_channel when it reasons through one.
The cutoff filter is the part that matters. first_seen_at on a relationship is
derived from article dates rather than processing time, so a historical proposal
only sees what the world had actually revealed by its own cutoff. Without that
this feature would quietly reintroduce the lookahead the evidence check exists to
prevent, and it is pinned by a test rather than left to review.
Relationships are explicitly background rather than evidence: predictions still
have to cite the article ids the story came from, and an instrument the articles
give no reason to care about is still not a prediction.
strategy_version moves to autonomy-2, because a prompt change this material
changes what a prediction means and the two populations should be comparable
later rather than silently blended.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
96% of rejected proposals name something untradable: indices (SPX, DXY, ^TNX),
fx (EURUSD, XAU/USD), futures (CL=F, BZ=F) and home listings (VOW3.DE, RHM.DE,
1211.HK, 688169.SS). The analysis behind those is usually sound, it is the ticker
that cannot be used, and nothing in the prompt ever said so. We were paying for
the call and discarding the result at validation.
The rules point the model at what the allowlist actually holds: US listings and
ADRs for foreign companies, and US listed ETFs as the tradable expression of an
index, currency, rate or commodity. Every symbol named in the rules was checked
against the live allowlist first, so VWAGY, BABA, TM, SONY, SPY, QQQ, GLD, USO,
UUP and TLT all genuinely resolve. It also forbids predicting SPY itself, which
is the benchmark and whose excess return is zero by construction.
Shared between the coordinator and replay prompts rather than written twice,
since a rule that drifts between the two lanes is worse than no rule.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The JSON shape in both prompts used real values as placeholders, and the model
was reading them as the answer:
instrument: 'NVDA' -> 397 of 611 predictions are NVDA (65%),
second place is LMT with 8
horizon_days: 10 -> 606 of 611 are horizon 10 (99.2%), out of
seven allowed horizons
direction: 'positive|negative' -> 502 of 611 are positive (82.2%)
event_type: 'stable_enum' -> the enum was never listed, so the model
invented one label per event, 201 distinct
values across 611 predictions
replayWorker had its own copy of the same prompt with the same values, which is
why both lanes show the identical skew (replay is 147/147 horizon 10, 138/147
NVDA).
Every placeholder is now a description of the field rather than a usable value,
with an explicit line saying not to copy them. event_type is validated against
the same closed family list the cohort key uses, so a label cannot mean one
thing in the prompt and another in calibration. Off-enum labels are salvaged
through the existing mapper when they are placeable and rejected when they are
not, so 'other' does not quietly become the bin again.
This does not by itself create edge. It means the next batch of predictions
measures the model's judgement instead of its willingness to copy an example.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Archive ingestion had been dead since 2026-08-02 because nothing in the
compose stack actually ran it. Everything downstream starved from there.
- add ingest + enrichment services. server.js only starts the scheduler when
DURIIN_RUN_SCHEDULER is not "false", and workers/index.js was not running at
all, so articles never got event_id/content/has_embedding and the coordinator
had nothing to lease.
- pass an explicit origin from coordinatorWorker. it was never passed, so
acceptProposal defaulted to 'live' and 464 historical backfill predictions
were recorded as live. that also meant verifyEvidence got a null cutoff and
skipped its date check entirely.
- coarsen cohortKey to event families + horizon buckets. 201 free text event
types produced 221 cohorts averaging 2.76 samples, so the n>=30 gate could
never be reached and everything abstained for the wrong reason.
- gate on cohort diversity, not just sample count. one ticker was roughly half
of all resolved outcomes, so a pure count gate was measuring one company.
unknown diversity abstains rather than passing.
- resolve the admin archive db explicitly and probe it. it relied on a
Dockerfile symlink, and without it better-sqlite3 quietly creates an empty
file and serves a phantom archive.
- clamp implausible future publication dates at ingest.
- pin the db backend to sqlite by default. compose hardcoded postgres "true",
which would have overridden the operator's own .env on the next redeploy and
pointed everything at a stale snapshot.
scripts/repair-autonomy-labels.js relabels the affected rows. it is dry run by
default and has not been applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb