Commit Graph
12 Commits
Author SHA1 Message Date
ImBenjiandClaude Opus 5 f464c94708 fix: drop untradable predictions instead of the whole proposal
One untradable ticker rejected everything alongside it. In a single day that was
93 proposals discarding 171 predictions, and 76 of those named something we could
trade perfectly well. They were lost because a sibling in the same response said
EURUSD.

Tradability is a filter, so it applies per prediction now. The untradable one is
dropped and logged, its siblings are kept, and the stored payload records what
was removed so the filtering is auditable rather than invisible.

Lookahead deliberately still rejects the entire proposal. Evidence that did not
exist at the proposal's own cutoff means the response is corrupt rather than
merely untradable, and keeping the rest of it would hide the one thing most worth
seeing. Both halves are pinned by tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-04 20:16:55 +01:00
ImBenjiandClaude Opus 5 ca69ad0e73 fix: stop three loops that retry forever, and let budget dead letters recover
The event outcome worker re-requested PSTG and GROQ on every poll for as long as
the process lived, because a fetch failure only logged and continued while the
prediction stayed pending. Ten requests every five minutes, indefinitely. Per
ticker backoff now doubles to an hour, so a symbol with no market data costs one
request an hour instead of one a minute. It also translates dotted tickers the
same way the autonomy worker does, which is why that helper moved into the
shared price module rather than being copied.

The gdelt loop had no pause on its error path at all, so once the api started
refusing connections it spun through failures continuously, burning cpu and
filling the log with the same stack. It has been doing that for days. Backs off
to half an hour now and resets on success.

isTransientCoordinatorFailure matched 408, 429 and 5xx but not a budget 402/403,
so the 380 jobs that dead-lettered during the exhausted quota window could never
come back on their own, including 55 live events. Budget failures are transient
in a way an ordinary auth failure is not, and a wrong key still dies permanently
because it says invalid or unauthorized rather than naming credits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-03 16:42:28 +01:00
ImBenjiandClaude Opus 5 859d0719b3 fix: stop discarding live predictions the moment they mature
Six live predictions were marked unresolvable, including MSFT twice, WMT and
ITW. Re-running the calculation against yahoo resolves all six, so they were
never unresolvable, they were scored before the market data existed and then
thrown away permanently.

The due-check counted calendar days while calculateOutcome finds the exit bar by
trading days. A friday horizon-1 prediction therefore looked due on saturday,
when monday's close cannot exist. calculateOutcome returned null and the worker
treated null as permanently dead. This hit short horizons hardest, which is
exactly the cohort that produces the first live evidence.

The sql filter stays loose because it cannot know about weekends, and trading day
arithmetic now decides what is genuinely ready. A null result waits for the
horizon to be properly past before anything is retired, and says so when it
finally gives up.

Separately, at horizon 1 the entry and exit lookups could land on the same bar
and produce an excess return of exactly zero, which was recorded as a real
outcome and scored as a directional miss. ITW and WMT both did this. A horizon
that has not elapsed is no longer a measurement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-03 16:28:57 +01:00
ImBenjiandClaude Opus 5 42fb9291b0 fix: unstick the content pipeline, the quota loop and the outcome retries
Three separate things had the pipeline frozen for 27 hours.

browserCrawler leaked page slots. context.newPage() sat outside the try, so a
throw or a hang there took the slot with it, and after maxConcurrentPages of
those every caller parked in acquirePageSlot forever. That is what it looked
like from outside: content workers alive, no logs, no progress, 13 chromium
renderers still up 10 hours after start. newPage is inside the try now, waiting
for a slot times out instead of blocking forever, and page.close() is raced so a
wedged renderer cant strand the slot on the way out either.

graphWorker had no backoff on quota failures. A blown OpenRouter monthly limit
returns an instant 403, so it retried as fast as the network allowed: 2356
failures in 20 minutes, drowning every other line in the log. Quota and auth
errors now pause resolution for 15 minutes and log once per window rather than
once per attempt.

The outcome worker retried unresolvable predictions forever. Yahoo writes class
shares with a dash, so BRK.B 404s every time, and a failed prediction stays open
and comes straight back on the next poll. Dots are translated to dashes, which
matters beyond this one name because the allowlist is full of dotted symbols,
and a prediction that fails five times is marked unresolvable instead of
spinning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-31 10:02:04 +01:00
ImBenjiandClaude Opus 5 b8b3987e35 fix: stop the coordinator copying its own prompt example
The JSON shape in both prompts used real values as placeholders, and the model
was reading them as the answer:

  instrument: 'NVDA'          -> 397 of 611 predictions are NVDA (65%),
                                 second place is LMT with 8
  horizon_days: 10            -> 606 of 611 are horizon 10 (99.2%), out of
                                 seven allowed horizons
  direction: 'positive|negative' -> 502 of 611 are positive (82.2%)
  event_type: 'stable_enum'   -> the enum was never listed, so the model
                                 invented one label per event, 201 distinct
                                 values across 611 predictions

replayWorker had its own copy of the same prompt with the same values, which is
why both lanes show the identical skew (replay is 147/147 horizon 10, 138/147
NVDA).

Every placeholder is now a description of the field rather than a usable value,
with an explicit line saying not to copy them. event_type is validated against
the same closed family list the cohort key uses, so a label cannot mean one
thing in the prompt and another in calibration. Off-enum labels are salvaged
through the existing mapper when they are placeable and rejected when they are
not, so 'other' does not quietly become the bin again.

This does not by itself create edge. It means the next batch of predictions
measures the model's judgement instead of its willingness to copy an example.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 23:29:11 +01:00
ImBenjiandClaude Opus 5 8f4d3b4ce9 fix: require live calibration to authorise a live order
Offline evidence can no longer authorise anything. createDecisions used to
prefer a live snapshot and fall back to the pooled historical one, so once the
live lane woke up a live prediction could have drawn a BUY off backfill data.
Backfill and replay are fine evidence that the pipeline works, they are not a
live track record.

No live snapshot now means ABSTAIN. The abstain says whether offline evidence
existed for that cohort, so "we have 60 offline samples but no live ones" stays
distinguishable from "we know nothing about this cohort".

Also split the health counter. It counted qualifying cohorts across every
source, which overstated how close we are to being able to trade now that only
live cohorts can authorise. It reports qualifying_live_cohorts and
qualifying_offline_cohorts separately, and applies the concentration cap it was
previously ignoring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 22:22:36 +01:00
ImBenjiandClaude Opus 5 6f1d1eee2d fix: restart the stalled autonomy pipeline and make calibration honest
Archive ingestion had been dead since 2026-08-02 because nothing in the
compose stack actually ran it. Everything downstream starved from there.

- add ingest + enrichment services. server.js only starts the scheduler when
  DURIIN_RUN_SCHEDULER is not "false", and workers/index.js was not running at
  all, so articles never got event_id/content/has_embedding and the coordinator
  had nothing to lease.
- pass an explicit origin from coordinatorWorker. it was never passed, so
  acceptProposal defaulted to 'live' and 464 historical backfill predictions
  were recorded as live. that also meant verifyEvidence got a null cutoff and
  skipped its date check entirely.
- coarsen cohortKey to event families + horizon buckets. 201 free text event
  types produced 221 cohorts averaging 2.76 samples, so the n>=30 gate could
  never be reached and everything abstained for the wrong reason.
- gate on cohort diversity, not just sample count. one ticker was roughly half
  of all resolved outcomes, so a pure count gate was measuring one company.
  unknown diversity abstains rather than passing.
- resolve the admin archive db explicitly and probe it. it relied on a
  Dockerfile symlink, and without it better-sqlite3 quietly creates an empty
  file and serves a phantom archive.
- clamp implausible future publication dates at ingest.
- pin the db backend to sqlite by default. compose hardcoded postgres "true",
  which would have overridden the operator's own .env on the next redeploy and
  pointed everything at a stale snapshot.

scripts/repair-autonomy-labels.js relabels the affected rows. it is dry run by
default and has not been applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 21:43:24 +01:00
ImBenji 4aa2367971 fix: calibrate from replay outcomes 2026-08-13 19:10:59 +01:00
ImBenji c9c2d0c8ef fix: recover transient coordinator dead letters 2026-08-13 19:02:41 +01:00
ImBenji 2c023c8962 fix: let replay recover from dead letter jobs 2026-08-08 23:32:41 +01:00
ImBenji 5877783862 feat: add isolated historical replay calibration 2026-08-04 22:00:11 +01:00
ImBenji c4028cc394 feat: add autonomous paper-trading and calibration pipeline 2026-08-03 14:03:27 +01:00