Removing the caps rather than tuning them. Every number I picked was a number I
invented, and 6000 was already tight enough to truncate a real replay article,
which is the failure that was called out when the cap first went in.
The coordinator now sends no max_tokens at all. The 402 handler supplies one only
when openrouter says the budget cannot cover an open ended request, so the
ceiling exists exactly when it has to and never otherwise. Verified unbounded is
accepted against the live key before making this the default.
The caps on the signal, augor, consolidation and graph workers are gone too. They
were added to work around an empty account, not because any of them ever produced
too much, and unlike the coordinator none of them detect truncation, so an
invented ceiling there risked silently corrupting company facts. A budget failure
in those is at least loud.
The 220 token cap in crawlerClassifier is left alone, it predates this and bounds
a genuinely tiny classification.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
signal, augor and consolidation had the same unbounded request as the
coordinator, so switching to a model with a 131k output window made all three
402 on every call while the coordinator itself was fine. Found them by grepping
for the endpoint rather than waiting for each one to surface in the logs.
Sized per worker rather than one global number, since these produce more than
the coordinator's small json object.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb