Content fetching stopped dead on 4 Sept at 17:06 and nobody noticed for three
days. In that window it managed about 328 articles and then produced nothing:
no error, no timeout, not a single line in the log. Meanwhile the live lane
starved, because an article needs content before it can be embedded, clustered
and handed to the coordinator, and 6596 articles arrived in 48 hours with zero
of them ready.
Every individual browser step already had a timeout. Acquiring the shared
session did not, and it is awaited while holding one of eight browser slots, so
a wedged chromium parks every slot permanently and nothing ever throws. The page
slot timeout added earlier never fired because it sits downstream of the thing
that was actually stuck. There is now an outer bound around the whole browser
path so the slot always comes back.
The workers also log the start and end of each round. The reason this took days
to find is that a healthy content worker and a completely wedged one looked
identical from outside, and that is worth fixing on its own.
Verified on the box first: outbound fetches return 200, the picker returns rows
in 4.8s, and fetchAndStoreContent stores a real article in 354ms. Every part
worked in isolation, which is what made the silence so misleading.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
The content backfill picker took 194 seconds per call. better-sqlite3 is
synchronous, so that blocked the whole ingest event loop, and with eight workers
each running it in a loop they serialised behind each other: about 26 minutes a
round. From outside it looked like a hang, and the gdelt loop went quiet at the
same time because it was stuck behind the same blocked loop.
Two causes. The picker tested `content IS NULL OR TRIM(content) = ''`, which
made sqlite read the content column, a 4GB blob, purely to decide which rows to
skip. content_status already records the same thing and agrees with the content
column on all 2.2M rows, so the test bought nothing. It also stopped any index
being usable.
Then there was no index matching the window function, so it built temp b-trees
over every unfetched row. idx_articles_pending_fetch is partial and column
ordered to match PARTITION BY source ORDER BY pub_date_effective DESC, id DESC.
The planner ignores it without stats, hence PRAGMA optimize.
Measured on production, same query, same 26k rows: 194.5s -> 1.03s.
Note for whoever reads this next: the playwright page slot leak fixed in 42fb929
was real but was not what froze the pipeline. This was.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb