Commit Graph
140 Commits
Author SHA1 Message Date
ImBenjiandClaude Opus 5 40f098129c perf: preload the ops module graph
Admin assets are no-store, so every load refetches all of them, and es module
imports are discovered one level at a time: main parses, then its imports are
found, then theirs. Preload hints turn that waterfall into one parallel burst.
The hrefs deliberately carry no version query because they have to byte-match
what the import specifiers resolve to, otherwise the browser fetches each module
twice instead of reusing the preload.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-01 20:26:43 +01:00
ImBenjiandClaude Opus 5 93108a27fe fix: convert style strings for react in the ops console
htm passes props through untouched and react rejects a style string with error
#62, so every view rendered an empty page. Converting once in the createElement
wrapper keeps plain css in the templates instead of style objects everywhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-01 20:20:27 +01:00
ImBenjiandClaude Opus 5 0a5b30ab63 feat: new ops console for the admin surface
The old admin was five separate html pages, so every navigation was a full
reload and the operational picture was scattered across all of them. Worth
saying: the api was never the problem, every endpoint answers in under 200ms.
It felt slow because of the architecture, not the backend.

This is a single page console. React and htm from a cdn, no bundler and no
babel-in-the-browser, because a runtime transpiler on every load is exactly the
slowness we are trying to get rid of. Hash routing, polling that keeps the last
good payload on screen instead of flashing a spinner, and stale responses are
dropped so a slow request cannot overwrite a newer one.

/admin/api/ops/overview answers the whole dashboard in one call rather than
making the browser fan out and stitch. It carries the things that actually
matter and were not visible anywhere before: when live evidence matures, why a
cohort does or does not clear the trade gate, which stage of the pipeline has
gone quiet, and what is sitting in dead letters.

Controls, all of which change production and all of which ask twice:
- requeue dead letters, which only ever moves dead_letter back to pending
- execution mode and a kill switch, now read from autonomy_settings on every
  poll instead of only from AUTONOMY_EXECUTION_MODE, so halting no longer needs
  a redeploy first
- run the reaction analysis and read its output

The d3 graph is framed rather than ported. It works, and rewriting it would risk
something valuable for nothing the operator can see.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-01 20:17:36 +01:00
ImBenjiandClaude Opus 5 2d92759ae9 perf: stop the content picker reading 4GB of article bodies
The content backfill picker took 194 seconds per call. better-sqlite3 is
synchronous, so that blocked the whole ingest event loop, and with eight workers
each running it in a loop they serialised behind each other: about 26 minutes a
round. From outside it looked like a hang, and the gdelt loop went quiet at the
same time because it was stuck behind the same blocked loop.

Two causes. The picker tested `content IS NULL OR TRIM(content) = ''`, which
made sqlite read the content column, a 4GB blob, purely to decide which rows to
skip. content_status already records the same thing and agrees with the content
column on all 2.2M rows, so the test bought nothing. It also stopped any index
being usable.

Then there was no index matching the window function, so it built temp b-trees
over every unfetched row. idx_articles_pending_fetch is partial and column
ordered to match PARTITION BY source ORDER BY pub_date_effective DESC, id DESC.
The planner ignores it without stats, hence PRAGMA optimize.

Measured on production, same query, same 26k rows: 194.5s -> 1.03s.

Note for whoever reads this next: the playwright page slot leak fixed in 42fb929
was real but was not what froze the pipeline. This was.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-09-01 15:03:03 +01:00
ImBenjiandClaude Opus 5 42fb9291b0 fix: unstick the content pipeline, the quota loop and the outcome retries
Three separate things had the pipeline frozen for 27 hours.

browserCrawler leaked page slots. context.newPage() sat outside the try, so a
throw or a hang there took the slot with it, and after maxConcurrentPages of
those every caller parked in acquirePageSlot forever. That is what it looked
like from outside: content workers alive, no logs, no progress, 13 chromium
renderers still up 10 hours after start. newPage is inside the try now, waiting
for a slot times out instead of blocking forever, and page.close() is raced so a
wedged renderer cant strand the slot on the way out either.

graphWorker had no backoff on quota failures. A blown OpenRouter monthly limit
returns an instant 403, so it retried as fast as the network allowed: 2356
failures in 20 minutes, drowning every other line in the log. Quota and auth
errors now pause resolution for 15 minutes and log once per window rather than
once per attempt.

The outcome worker retried unresolvable predictions forever. Yahoo writes class
shares with a dash, so BRK.B 404s every time, and a failed prediction stays open
and comes straight back on the next poll. Dots are translated to dashes, which
matters beyond this one name because the allowlist is full of dotted symbols,
and a prediction that fails five times is marked unresolvable instead of
spinning.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-31 10:02:04 +01:00
ImBenjiandClaude Opus 5 b8b3987e35 fix: stop the coordinator copying its own prompt example
The JSON shape in both prompts used real values as placeholders, and the model
was reading them as the answer:

  instrument: 'NVDA'          -> 397 of 611 predictions are NVDA (65%),
                                 second place is LMT with 8
  horizon_days: 10            -> 606 of 611 are horizon 10 (99.2%), out of
                                 seven allowed horizons
  direction: 'positive|negative' -> 502 of 611 are positive (82.2%)
  event_type: 'stable_enum'   -> the enum was never listed, so the model
                                 invented one label per event, 201 distinct
                                 values across 611 predictions

replayWorker had its own copy of the same prompt with the same values, which is
why both lanes show the identical skew (replay is 147/147 horizon 10, 138/147
NVDA).

Every placeholder is now a description of the field rather than a usable value,
with an explicit line saying not to copy them. event_type is validated against
the same closed family list the cohort key uses, so a label cannot mean one
thing in the prompt and another in calibration. Off-enum labels are salvaged
through the existing mapper when they are placeable and rejected when they are
not, so 'other' does not quietly become the bin again.

This does not by itself create edge. It means the next batch of predictions
measures the model's judgement instead of its willingness to copy an example.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 23:29:11 +01:00
ImBenjiandClaude Opus 5 c4650fe45c feat: measure whether the initial market reaction conditions anything
The coordinator is scored on excess return vs SPY starting at the information
cutoff, so the announcement move sits outside the scored window. That move is
the best documented conditioner for post event drift and we were discarding it.
This measures whether keeping it would buy us anything, before any of it gets
wired into the cohort key or the prompt.

Reaction is measured from the last close before the event's first article up to
the outcome's own entry price, so the reaction and forward windows touch but
never overlap. Read only, and it caches price history so it can be re-run cheaply
as more outcomes mature.

T1 and T2 are pre-registered in the header because sweeping buckets over 611
outcomes that are half one ticker will always turn up something.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 23:02:06 +01:00
ImBenjiandClaude Opus 5 8f4d3b4ce9 fix: require live calibration to authorise a live order
Offline evidence can no longer authorise anything. createDecisions used to
prefer a live snapshot and fall back to the pooled historical one, so once the
live lane woke up a live prediction could have drawn a BUY off backfill data.
Backfill and replay are fine evidence that the pipeline works, they are not a
live track record.

No live snapshot now means ABSTAIN. The abstain says whether offline evidence
existed for that cohort, so "we have 60 offline samples but no live ones" stays
distinguishable from "we know nothing about this cohort".

Also split the health counter. It counted qualifying cohorts across every
source, which overstated how close we are to being able to trade now that only
live cohorts can authorise. It reports qualifying_live_cohorts and
qualifying_offline_cohorts separately, and applies the concentration cap it was
previously ignoring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 22:22:36 +01:00
ImBenjiandClaude Opus 5 6f1d1eee2d fix: restart the stalled autonomy pipeline and make calibration honest
Archive ingestion had been dead since 2026-08-02 because nothing in the
compose stack actually ran it. Everything downstream starved from there.

- add ingest + enrichment services. server.js only starts the scheduler when
  DURIIN_RUN_SCHEDULER is not "false", and workers/index.js was not running at
  all, so articles never got event_id/content/has_embedding and the coordinator
  had nothing to lease.
- pass an explicit origin from coordinatorWorker. it was never passed, so
  acceptProposal defaulted to 'live' and 464 historical backfill predictions
  were recorded as live. that also meant verifyEvidence got a null cutoff and
  skipped its date check entirely.
- coarsen cohortKey to event families + horizon buckets. 201 free text event
  types produced 221 cohorts averaging 2.76 samples, so the n>=30 gate could
  never be reached and everything abstained for the wrong reason.
- gate on cohort diversity, not just sample count. one ticker was roughly half
  of all resolved outcomes, so a pure count gate was measuring one company.
  unknown diversity abstains rather than passing.
- resolve the admin archive db explicitly and probe it. it relied on a
  Dockerfile symlink, and without it better-sqlite3 quietly creates an empty
  file and serves a phantom archive.
- clamp implausible future publication dates at ingest.
- pin the db backend to sqlite by default. compose hardcoded postgres "true",
  which would have overridden the operator's own .env on the next redeploy and
  pointed everything at a stale snapshot.

scripts/repair-autonomy-labels.js relabels the affected rows. it is dry run by
default and has not been applied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
2026-08-29 21:43:24 +01:00
ImBenji d778a02bfb fix: make postgres the default runtime service 2026-08-19 09:12:44 +01:00
ImBenji 9eb59443bd fix: support postgres pragma table info 2026-08-19 08:42:58 +01:00
ImBenji 8288bbc028 fix: stop eager admin page prefetch 2026-08-19 08:38:12 +01:00
ImBenji 7ff69ad99d fix: use async postgres for autonomy http routes 2026-08-17 14:10:29 +01:00
ImBenji aea0ed437e fix: isolate autonomy postgres url 2026-08-17 14:05:31 +01:00
ImBenji 51649fbf5f fix: pin postgres transaction clients 2026-08-17 14:02:59 +01:00
ImBenji 4c24b37af9 fix: keep postgres archive time lookup indexed 2026-08-17 13:59:13 +01:00
ImBenji e239a5d41e fix: cast postgres archive timestamps 2026-08-17 13:56:13 +01:00
ImBenji f0b598a3b8 feat: support postgres autonomy runtime 2026-08-17 13:43:50 +01:00
ImBenji 4aa2367971 fix: calibrate from replay outcomes 2026-08-13 19:10:59 +01:00
ImBenji c9c2d0c8ef fix: recover transient coordinator dead letters 2026-08-13 19:02:41 +01:00
ImBenji 2c023c8962 fix: let replay recover from dead letter jobs 2026-08-08 23:32:41 +01:00
ImBenji 0344d5ca97 Revert "style: reduce modal title scale"
This reverts commit 2d5c0280de.
2026-08-04 22:39:50 +01:00
ImBenji 2d5c0280de style: reduce modal title scale 2026-08-04 22:38:37 +01:00
ImBenji 0047adcb6d fix: scrub invalid nul bytes during postgres migration 2026-08-04 22:35:50 +01:00
ImBenji ffc85719ff feat: add bounded postgres data migration stack 2026-08-04 22:09:23 +01:00
ImBenji 5877783862 feat: add isolated historical replay calibration 2026-08-04 22:00:11 +01:00
ImBenji be18eb77f0 fix: cache background statistics response 2026-08-04 21:51:18 +01:00
ImBenji b3a55fac08 perf: source statistics from the catalog 2026-08-04 21:50:02 +01:00
ImBenji 6e666838ba perf: keep source aggregation off interactive requests 2026-08-04 21:48:44 +01:00
ImBenji d020fa3811 perf: avoid full archive scans in statistics 2026-08-04 21:46:56 +01:00
ImBenji f5f092cfb0 perf: compute archive stats off the API thread 2026-08-04 21:45:17 +01:00
ImBenji 002b7b320c perf: remove blocking archive counts from navigation 2026-08-04 21:42:23 +01:00
ImBenji 2020a0fde2 fix: keep archive indexing off API startup 2026-08-04 21:39:20 +01:00
ImBenji bd162606b6 perf: make admin navigation responsive 2026-08-04 21:36:39 +01:00
ImBenji 5680ee8a44 fix: prevent stale admin UI assets 2026-08-04 21:31:55 +01:00
ImBenji b9ee10a83a fix: serialize autonomy job leases 2026-08-04 21:07:42 +01:00
ImBenji 5037e1e192 feat: overhaul Duriin autonomy console 2026-08-04 21:05:47 +01:00
ImBenji 0724f5dc36 fix: run autonomy workers by default 2026-08-03 14:43:30 +01:00
ImBenji c4028cc394 feat: add autonomous paper-trading and calibration pipeline 2026-08-03 14:03:27 +01:00
ImBenji 5a9a2e4c6d refactor: load environment variables from .env file and update openRouter configuration 2026-04-27 19:08:54 +01:00
ImBenji 80d545fa6a refactor: load environment variables from .env file and update openRouter configuration 2026-04-27 19:00:32 +01:00
ImBenji 1ec273a72b refactor: load environment variables from .env file and update openRouter configuration 2026-04-27 18:39:35 +01:00
ImBenji 82abe0bcb3 refactor: load environment variables from .env file and update openRouter configuration 2026-04-27 18:02:20 +01:00
ImBenji 7ceaaf2401 refactor: load environment variables from .env file and update openRouter configuration 2026-04-27 17:43:48 +01:00
ImBenji 80236b9396 refactor: add env_file configuration to docker-compose.yml for environment variable management 2026-04-27 17:40:15 +01:00
ImBenji 1d77266cb2 refactor: enhance database URL handling in config.js for improved PostgreSQL connection setup 2026-04-27 17:38:41 +01:00
ImBenji 82df9da814 refactor: adjust directory navigation in rebuild-api.sh for improved script execution 2026-04-27 17:34:18 +01:00
ImBenji 94281b81e2 refactor: update worker commands and add new scripts for API rebuilding and queue feeding 2026-04-27 17:31:45 +01:00
ImBenji f00b1a9640 refactor: update worker commands and add new scripts for API rebuilding and queue feeding 2026-04-27 17:30:52 +01:00
ImBenji 04966fac55 refactor: update worker commands and add new scripts for API rebuilding and queue feeding 2026-04-27 15:03:15 +01:00