perf: stop the content picker reading 4GB of article bodies

The content backfill picker took 194 seconds per call. better-sqlite3 is
synchronous, so that blocked the whole ingest event loop, and with eight workers
each running it in a loop they serialised behind each other: about 26 minutes a
round. From outside it looked like a hang, and the gdelt loop went quiet at the
same time because it was stuck behind the same blocked loop.

Two causes. The picker tested `content IS NULL OR TRIM(content) = ''`, which
made sqlite read the content column, a 4GB blob, purely to decide which rows to
skip. content_status already records the same thing and agrees with the content
column on all 2.2M rows, so the test bought nothing. It also stopped any index
being usable.

Then there was no index matching the window function, so it built temp b-trees
over every unfetched row. idx_articles_pending_fetch is partial and column
ordered to match PARTITION BY source ORDER BY pub_date_effective DESC, id DESC.
The planner ignores it without stats, hence PRAGMA optimize.

Measured on production, same query, same 26k rows: 194.5s -> 1.03s.

Note for whoever reads this next: the playwright page slot leak fixed in 42fb929
was real but was not what froze the pipeline. This was.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
This commit is contained in:
ImBenji
2026-09-01 15:03:03 +01:00
co-authored by Claude Opus 5
parent 42fb9291b0
commit 2d92759ae9
2 changed files with 24 additions and 4 deletions
+6 -4
View File
@@ -68,8 +68,11 @@ const selectPartitionedArticlesMissingContent = db.prepare(`
SELECT id, url, title, description, source, pub_date_effective,
ROW_NUMBER() OVER (PARTITION BY source ORDER BY pub_date_effective DESC, id DESC) AS rn
FROM articles
WHERE (content IS NULL OR TRIM(content) = '')
AND (content_status IS NULL OR content_status = 'pending')
-- content_status is the authority on whether a row has been fetched, and it
-- agrees with the content column on all 2.2M rows. Testing TRIM(content) here
-- as well meant reading a 4GB blob column just to find out which rows to skip,
-- and it stopped the partial index below being usable at all.
WHERE (content_status IS NULL OR content_status = 'pending')
AND (content_retry_after IS NULL OR content_retry_after <= datetime('now'))
AND (id % ?) = ?
)
@@ -409,8 +412,7 @@ async function runBackfillWorker({ workerIndex, workerCount, perSource, batchSiz
function hasPendingContent() {
return Boolean(db.prepare(`
SELECT 1 FROM articles
WHERE (content IS NULL OR TRIM(content) = '')
AND (content_status IS NULL OR content_status = 'pending')
WHERE (content_status IS NULL OR content_status = 'pending')
AND (content_retry_after IS NULL OR content_retry_after <= datetime('now'))
LIMIT 1
`).get());
+18
View File
@@ -54,8 +54,26 @@ db.exec(`
AND content != ''
AND is_index_page = 0
AND has_embedding = 1;
-- The content backfill picker partitions by source and orders by pub date, and
-- without this it built two temp b-trees over every unfetched row: 194 seconds
-- per call on a 2.2M row archive, synchronously, which froze the whole process.
-- Column order matches PARTITION BY source ORDER BY pub_date_effective DESC, id DESC
-- so the window function can just walk it.
CREATE INDEX IF NOT EXISTS idx_articles_pending_fetch
ON articles(source, pub_date_effective DESC, id DESC)
WHERE content_status IS NULL OR content_status = 'pending';
`);
// Without stats the planner ignores the partial index above and falls back to the
// content_status index plus a temp b-tree, which is roughly 70% slower. optimize
// only re-analyses what has actually drifted, so this is cheap after the first run.
try {
db.exec('PRAGMA optimize;');
} catch (error) {
console.error('[db] PRAGMA optimize failed, query plans may be stale:', error.message);
}
db.exec(`
CREATE TABLE IF NOT EXISTS article_embedding_store (
article_id INTEGER NOT NULL,