The brief tells the model an empty predictions array is always available, so
run 2 may answer fewer articles than run 1 did. Writing down now, while there
are two run-2 proposals on the board, that selectivity gets reported as a
result rather than quietly treated as a smaller sample.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
leaseNextJob hands back any pending replay_article job, it has no idea about
runs, and the worker was attributing whatever came back to whichever run was
active. One recovered dead letter from run 1 would have been stamped with run
2's id, given run 2's feedback brief, and dragged run 2's cursor to wherever
that old article sits in the archive. A pinned run would then decide its set was
finished after a couple of articles. There are 139 dead letters and they are
built to recover, so this was not hypothetical.
The idempotency key already says which run enqueued the job. Ask it.
Also: refuse to inherit the parent's model label when starting a run. Inheriting
is exactly how run 1 came to be labelled qwen for predictions deepseek made.
The split moves to the replay container's actual restart time rather than the
commit timestamp five minutes later. Verified the running container really does
have the instrument rules, the de-anchoring and the enum before trusting it as
the boundary. It makes no difference to the partition, there are no replay
predictions at all between 15:57 and midnight that day, but the boundary should
be the thing that actually changed the prompt.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Nothing in the pipeline has ever fed an outcome back to the thing that makes
predictions. Calibration reads autonomy_outcomes, but calibration only gates
whether to act on a prediction, never what the prediction is. So the only thing
that has ever changed this system's output is a human editing the prompt.
score-replay-runs.js asks "compared to what". The answer is not flattering:
over 2,114 scored replay predictions the system is right 50.05% of the time
while answering "negative" to every one of the same bars scores 54.45%. It is
4.4 points below a constant, z=-4.06. The whole deficit is the prior. It says
positive on 65% of calls when 45.5% of bars beat SPY, a 19 point skew. Its
discrimination, P(up|positive) minus P(up|negative), is +3.2 points with
p=0.15, so the direction it picks is weakly informative and completely buried
by how often it defaults to positive.
The first version of that script compared each direction group's accuracy to
"always that direction" on the same rows, which is an identity and tests
nothing. T3 replaces it with the two proportion test that actually asks whether
the choice of direction carries information.
build-feedback-brief.js turns a run's scored outcomes into a memo the next run
reads before predicting. Generated from the data, not written by hand, or it is
just me editing the prompt again with extra steps.
Replay can now be pinned to an explicit article set, which is what makes two
runs comparable at all. Comparing two calendar windows of one run compares two
market regimes: the epochs in run 1 line up exactly with article vintage, E0 is
late 2024 and E2 is 2026, so nothing could be attributed. A new run also
inherits its parent's watermark instead of recomputing it from today, which
silently guaranteed a different archive slice every time.
prompt_version never moved across four material prompt changes, so every
proposal on record claims to come from the first prompt. coordinator-2 and
replay-coordinator-2.
docs/replay-run-2-preregistration.md fixes the bar before the run exists,
including which result counts as learning and which is only calibration.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Live outcomes have been stuck at 4 while 35 matured live predictions sat
unscored, the oldest four days past its horizon. Running those three by hand
resolves them in 200-400ms each, so the data was there and the maths was fine.
The worker itself was wedged: up three days, last log 6 Sept, processing
nothing.
req.setTimeout only covers socket inactivity. A response that opens and then
stalls leaves the promise pending forever, and with it the entire loop, because
the fetch is awaited inline. There is now a hard bound around it.
This is the third time an unbounded await inside a long lived loop has silently
stopped a worker: the browser session in content, the content round itself, and
now market data. In every case the container stayed up, nothing threw, and
nothing was logged, which is the worst possible failure shape. So the worker also
announces what it is about to score, because an idle worker and a dead one
should not look identical from outside.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Content fetching stopped dead on 4 Sept at 17:06 and nobody noticed for three
days. In that window it managed about 328 articles and then produced nothing:
no error, no timeout, not a single line in the log. Meanwhile the live lane
starved, because an article needs content before it can be embedded, clustered
and handed to the coordinator, and 6596 articles arrived in 48 hours with zero
of them ready.
Every individual browser step already had a timeout. Acquiring the shared
session did not, and it is awaited while holding one of eight browser slots, so
a wedged chromium parks every slot permanently and nothing ever throws. The page
slot timeout added earlier never fired because it sits downstream of the thing
that was actually stuck. There is now an outer bound around the whole browser
path so the slot always comes back.
The workers also log the start and end of each round. The reason this took days
to find is that a healthy content worker and a completely wedged one looked
identical from outside, and that is worth fixing on its own.
Verified on the box first: outbound fetches return 200, the picker returns rows
in 4.8s, and fetchAndStoreContent stores a real article in 354ms. Every part
worked in isolation, which is what made the silence so misleading.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The graph has been built for months and fed nothing but a dashboard. Nothing in
the autonomy pipeline has ever read an edge: not the coordinator, not
calibration, not execution. Its only downstream consumer, trade_signals, last
produced anything in April. It was roughly 70% of the llm bill and informed no
prediction, decision or order.
Relationships are the one piece of context a per-event coordinator genuinely
cannot derive from its own articles, because "this company supplies that one" is
knowledge about companies rather than about this event. So the coordinator now
receives the relationships of the companies the event is about, and is told to
name the relationship in causal_channel when it reasons through one.
The cutoff filter is the part that matters. first_seen_at on a relationship is
derived from article dates rather than processing time, so a historical proposal
only sees what the world had actually revealed by its own cutoff. Without that
this feature would quietly reintroduce the lookahead the evidence check exists to
prevent, and it is pinned by a test rather than left to review.
Relationships are explicitly background rather than evidence: predictions still
have to cite the article ids the story came from, and an instrument the articles
give no reason to care about is still not a prediction.
strategy_version moves to autonomy-2, because a prompt change this material
changes what a prediction means and the two populations should be comparable
later rather than silently blended.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The env var has been set to google/gemma-4-31b-it all along and nothing ever
mapped it onto openRouter.cheapModel, so graphWorker's fallback chain silently
used the main model instead. Graph entity resolution is the highest volume llm
call in the system and its entire job is to reply with the number of a match.
Measured on the live key, same prompt:
deepseek v4 flash 7564 completion tokens, 7558 of them reasoning $0.0013708
gemma-4-31b-it 14 completion tokens, 0 reasoning $0.0000121
113x. Gemma is actually the more expensive model per token, which is why this
was worth measuring rather than reasoning about prices: the cost is not the
price of the tokens, it is a reasoning model spending seven thousand tokens
thinking about a multiple choice question.
This also explains the reasoning tokens dominating the usage dashboard, and why
the daily spend roughly doubled today rather than yesterday. Removing the token
ceiling let a trivial prompt reason without bound. The ceiling was never the
right control for that, the model choice is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
One untradable ticker rejected everything alongside it. In a single day that was
93 proposals discarding 171 predictions, and 76 of those named something we could
trade perfectly well. They were lost because a sibling in the same response said
EURUSD.
Tradability is a filter, so it applies per prediction now. The untradable one is
dropped and logged, its siblings are kept, and the stored payload records what
was removed so the filtering is auditable rather than invisible.
Lookahead deliberately still rejects the entire proposal. Evidence that did not
exist at the proposal's own cutoff means the response is corrupt rather than
merely untradable, and keeping the rest of it would hide the one thing most worth
seeing. Both halves are pinned by tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
96% of rejected proposals name something untradable: indices (SPX, DXY, ^TNX),
fx (EURUSD, XAU/USD), futures (CL=F, BZ=F) and home listings (VOW3.DE, RHM.DE,
1211.HK, 688169.SS). The analysis behind those is usually sound, it is the ticker
that cannot be used, and nothing in the prompt ever said so. We were paying for
the call and discarding the result at validation.
The rules point the model at what the allowlist actually holds: US listings and
ADRs for foreign companies, and US listed ETFs as the tradable expression of an
index, currency, rate or commodity. Every symbol named in the rules was checked
against the live allowlist first, so VWAGY, BABA, TM, SONY, SPY, QQQ, GLD, USO,
UUP and TLT all genuinely resolve. It also forbids predicting SPY itself, which
is the benchmark and whose excess return is zero by construction.
Shared between the coordinator and replay prompts rather than written twice,
since a rule that drifts between the two lanes is worse than no rule.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Removing the caps rather than tuning them. Every number I picked was a number I
invented, and 6000 was already tight enough to truncate a real replay article,
which is the failure that was called out when the cap first went in.
The coordinator now sends no max_tokens at all. The 402 handler supplies one only
when openrouter says the budget cannot cover an open ended request, so the
ceiling exists exactly when it has to and never otherwise. Verified unbounded is
accepted against the live key before making this the default.
The caps on the signal, augor, consolidation and graph workers are gone too. They
were added to work around an empty account, not because any of them ever produced
too much, and unlike the coordinator none of them detect truncation, so an
invented ceiling there risked silently corrupting company facts. A budget failure
in those is at least loud.
The 220 token cap in crawlerClassifier is left alone, it predates this and bounds
a genuinely tiny classification.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
A replay article dead-lettered with 'coordinator response was truncated by
max_tokens'. The cap was too tight, which is exactly the failure that was
predicted when it went in.
Raising it to 32000 costs nothing. The cap is there to satisfy openrouter's
affordability check, not to ration tokens, and billing is on tokens used rather
than reserved. The 402 handler already walks the ceiling down automatically when
the budget cannot cover the reservation, so a high default is effectively
uncapped while funded and degrades by itself when not.
Verified against the live key at 64000 before picking 32000, so there is real
headroom rather than a guess.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The event outcome worker re-requested PSTG and GROQ on every poll for as long as
the process lived, because a fetch failure only logged and continued while the
prediction stayed pending. Ten requests every five minutes, indefinitely. Per
ticker backoff now doubles to an hour, so a symbol with no market data costs one
request an hour instead of one a minute. It also translates dotted tickers the
same way the autonomy worker does, which is why that helper moved into the
shared price module rather than being copied.
The gdelt loop had no pause on its error path at all, so once the api started
refusing connections it spun through failures continuously, burning cpu and
filling the log with the same stack. It has been doing that for days. Backs off
to half an hour now and resets on success.
isTransientCoordinatorFailure matched 408, 429 and 5xx but not a budget 402/403,
so the 380 jobs that dead-lettered during the exhausted quota window could never
come back on their own, including 55 live events. Budget failures are transient
in a way an ordinary auth failure is not, and a wrong key still dies permanently
because it says invalid or unauthorized rather than naming credits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Six live predictions were marked unresolvable, including MSFT twice, WMT and
ITW. Re-running the calculation against yahoo resolves all six, so they were
never unresolvable, they were scored before the market data existed and then
thrown away permanently.
The due-check counted calendar days while calculateOutcome finds the exit bar by
trading days. A friday horizon-1 prediction therefore looked due on saturday,
when monday's close cannot exist. calculateOutcome returned null and the worker
treated null as permanently dead. This hit short horizons hardest, which is
exactly the cohort that produces the first live evidence.
The sql filter stays loose because it cannot know about weekends, and trading day
arithmetic now decides what is genuinely ready. A null result waits for the
horizon to be properly past before anything is retired, and says so when it
finally gives up.
Separately, at horizon 1 the entry and exit lookups could land on the same bar
and produce an excess return of exactly zero, which was recorded as a real
outcome and scored as a directional miss. ITW and WMT both did this. A horizon
that has not elapsed is no longer a measurement.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Two things wrong with the fixed cap I added earlier.
It was guessed rather than measured. The largest proposal this has ever produced
was about 1,258 tokens carrying 12 predictions and the average is around 50, so
8000 was arbitrary, and worse, it was above what the key could afford by the
time it deployed. The ceiling openrouter will accept shrinks as the balance
depletes: 15,666 earlier today, 3,921 an hour later.
So the cap is now adaptive. A 402 names the ceiling, and we retry once just under
it, downwards only. A shrinking budget shortens the allowed answer instead of
stopping the pipeline dead. Worth being clear that removing the cap is not an
option on a limited key, an unbounded request is refused outright and produces
no output at all rather than a truncated one.
And truncation is no longer silent. finish_reason length now throws instead of
handing a half written response to extractJson, which could occasionally parse a
partial object and quietly drop predictions. That was the real risk in capping.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
signal, augor and consolidation had the same unbounded request as the
coordinator, so switching to a model with a 131k output window made all three
402 on every call while the coordinator itself was fine. Found them by grepping
for the endpoint rather than waiting for each one to surface in the logs.
Sized per worker rather than one global number, since these produce more than
the coordinator's small json object.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Neither the coordinator nor the graph resolver set max_tokens, so openrouter
reserved the model's entire output window against the key's remaining budget and
returned 402 before running anything: 131k tokens reserved to produce a few
hundred. It never showed up with the old model because its output window is
small enough to fit under the limit.
Both are now bounded and overridable by env. The headroom is deliberate, the
newer reasoning models spend completion tokens thinking before they answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Measured on the deployed console: the page shell and every module land in 39ms,
so the 1.4s was entirely /admin/api/ops/overview, fetched twice.
Twice because the sidebar and the overview view each called usePoll on the same
url, each with its own timer. usePoll now keeps one store per url, so any number
of subscribers share a single request and a single interval, and a request
already on the wire is joined rather than duplicated.
The endpoint itself was dominated by the jobs rollup: a full scan of
autonomy_jobs, 521k rows and growing about nine thousand a day. A covering index
on (job_type, lane, status, created_at) takes it from 548ms to 110ms.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Admin assets are no-store, so every load refetches all of them, and es module
imports are discovered one level at a time: main parses, then its imports are
found, then theirs. Preload hints turn that waterfall into one parallel burst.
The hrefs deliberately carry no version query because they have to byte-match
what the import specifiers resolve to, otherwise the browser fetches each module
twice instead of reusing the preload.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
htm passes props through untouched and react rejects a style string with error
#62, so every view rendered an empty page. Converting once in the createElement
wrapper keeps plain css in the templates instead of style objects everywhere.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The old admin was five separate html pages, so every navigation was a full
reload and the operational picture was scattered across all of them. Worth
saying: the api was never the problem, every endpoint answers in under 200ms.
It felt slow because of the architecture, not the backend.
This is a single page console. React and htm from a cdn, no bundler and no
babel-in-the-browser, because a runtime transpiler on every load is exactly the
slowness we are trying to get rid of. Hash routing, polling that keeps the last
good payload on screen instead of flashing a spinner, and stale responses are
dropped so a slow request cannot overwrite a newer one.
/admin/api/ops/overview answers the whole dashboard in one call rather than
making the browser fan out and stitch. It carries the things that actually
matter and were not visible anywhere before: when live evidence matures, why a
cohort does or does not clear the trade gate, which stage of the pipeline has
gone quiet, and what is sitting in dead letters.
Controls, all of which change production and all of which ask twice:
- requeue dead letters, which only ever moves dead_letter back to pending
- execution mode and a kill switch, now read from autonomy_settings on every
poll instead of only from AUTONOMY_EXECUTION_MODE, so halting no longer needs
a redeploy first
- run the reaction analysis and read its output
The d3 graph is framed rather than ported. It works, and rewriting it would risk
something valuable for nothing the operator can see.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The content backfill picker took 194 seconds per call. better-sqlite3 is
synchronous, so that blocked the whole ingest event loop, and with eight workers
each running it in a loop they serialised behind each other: about 26 minutes a
round. From outside it looked like a hang, and the gdelt loop went quiet at the
same time because it was stuck behind the same blocked loop.
Two causes. The picker tested `content IS NULL OR TRIM(content) = ''`, which
made sqlite read the content column, a 4GB blob, purely to decide which rows to
skip. content_status already records the same thing and agrees with the content
column on all 2.2M rows, so the test bought nothing. It also stopped any index
being usable.
Then there was no index matching the window function, so it built temp b-trees
over every unfetched row. idx_articles_pending_fetch is partial and column
ordered to match PARTITION BY source ORDER BY pub_date_effective DESC, id DESC.
The planner ignores it without stats, hence PRAGMA optimize.
Measured on production, same query, same 26k rows: 194.5s -> 1.03s.
Note for whoever reads this next: the playwright page slot leak fixed in 42fb929
was real but was not what froze the pipeline. This was.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Three separate things had the pipeline frozen for 27 hours.
browserCrawler leaked page slots. context.newPage() sat outside the try, so a
throw or a hang there took the slot with it, and after maxConcurrentPages of
those every caller parked in acquirePageSlot forever. That is what it looked
like from outside: content workers alive, no logs, no progress, 13 chromium
renderers still up 10 hours after start. newPage is inside the try now, waiting
for a slot times out instead of blocking forever, and page.close() is raced so a
wedged renderer cant strand the slot on the way out either.
graphWorker had no backoff on quota failures. A blown OpenRouter monthly limit
returns an instant 403, so it retried as fast as the network allowed: 2356
failures in 20 minutes, drowning every other line in the log. Quota and auth
errors now pause resolution for 15 minutes and log once per window rather than
once per attempt.
The outcome worker retried unresolvable predictions forever. Yahoo writes class
shares with a dash, so BRK.B 404s every time, and a failed prediction stays open
and comes straight back on the next poll. Dots are translated to dashes, which
matters beyond this one name because the allowlist is full of dotted symbols,
and a prediction that fails five times is marked unresolvable instead of
spinning.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The JSON shape in both prompts used real values as placeholders, and the model
was reading them as the answer:
instrument: 'NVDA' -> 397 of 611 predictions are NVDA (65%),
second place is LMT with 8
horizon_days: 10 -> 606 of 611 are horizon 10 (99.2%), out of
seven allowed horizons
direction: 'positive|negative' -> 502 of 611 are positive (82.2%)
event_type: 'stable_enum' -> the enum was never listed, so the model
invented one label per event, 201 distinct
values across 611 predictions
replayWorker had its own copy of the same prompt with the same values, which is
why both lanes show the identical skew (replay is 147/147 horizon 10, 138/147
NVDA).
Every placeholder is now a description of the field rather than a usable value,
with an explicit line saying not to copy them. event_type is validated against
the same closed family list the cohort key uses, so a label cannot mean one
thing in the prompt and another in calibration. Off-enum labels are salvaged
through the existing mapper when they are placeable and rejected when they are
not, so 'other' does not quietly become the bin again.
This does not by itself create edge. It means the next batch of predictions
measures the model's judgement instead of its willingness to copy an example.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The coordinator is scored on excess return vs SPY starting at the information
cutoff, so the announcement move sits outside the scored window. That move is
the best documented conditioner for post event drift and we were discarding it.
This measures whether keeping it would buy us anything, before any of it gets
wired into the cohort key or the prompt.
Reaction is measured from the last close before the event's first article up to
the outcome's own entry price, so the reaction and forward windows touch but
never overlap. Read only, and it caches price history so it can be re-run cheaply
as more outcomes mature.
T1 and T2 are pre-registered in the header because sweeping buckets over 611
outcomes that are half one ticker will always turn up something.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Offline evidence can no longer authorise anything. createDecisions used to
prefer a live snapshot and fall back to the pooled historical one, so once the
live lane woke up a live prediction could have drawn a BUY off backfill data.
Backfill and replay are fine evidence that the pipeline works, they are not a
live track record.
No live snapshot now means ABSTAIN. The abstain says whether offline evidence
existed for that cohort, so "we have 60 offline samples but no live ones" stays
distinguishable from "we know nothing about this cohort".
Also split the health counter. It counted qualifying cohorts across every
source, which overstated how close we are to being able to trade now that only
live cohorts can authorise. It reports qualifying_live_cohorts and
qualifying_offline_cohorts separately, and applies the concentration cap it was
previously ignoring.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
Archive ingestion had been dead since 2026-08-02 because nothing in the
compose stack actually ran it. Everything downstream starved from there.
- add ingest + enrichment services. server.js only starts the scheduler when
DURIIN_RUN_SCHEDULER is not "false", and workers/index.js was not running at
all, so articles never got event_id/content/has_embedding and the coordinator
had nothing to lease.
- pass an explicit origin from coordinatorWorker. it was never passed, so
acceptProposal defaulted to 'live' and 464 historical backfill predictions
were recorded as live. that also meant verifyEvidence got a null cutoff and
skipped its date check entirely.
- coarsen cohortKey to event families + horizon buckets. 201 free text event
types produced 221 cohorts averaging 2.76 samples, so the n>=30 gate could
never be reached and everything abstained for the wrong reason.
- gate on cohort diversity, not just sample count. one ticker was roughly half
of all resolved outcomes, so a pure count gate was measuring one company.
unknown diversity abstains rather than passing.
- resolve the admin archive db explicitly and probe it. it relied on a
Dockerfile symlink, and without it better-sqlite3 quietly creates an empty
file and serves a phantom archive.
- clamp implausible future publication dates at ingest.
- pin the db backend to sqlite by default. compose hardcoded postgres "true",
which would have overridden the operator's own .env on the next redeploy and
pointed everything at a stale snapshot.
scripts/repair-autonomy-labels.js relabels the affected rows. it is dry run by
default and has not been applied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb