fix: bound max_tokens on the openrouter calls
Neither the coordinator nor the graph resolver set max_tokens, so openrouter reserved the model's entire output window against the key's remaining budget and returned 402 before running anything: 131k tokens reserved to produce a few hundred. It never showed up with the old model because its output window is small enough to fit under the limit. Both are now bounded and overridable by env. The headroom is deliberate, the newer reasoning models spend completion tokens thinking before they answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb
This commit is contained in:
@@ -23,6 +23,11 @@ async function callCoordinator(config, prompt) {
|
||||
body: JSON.stringify({
|
||||
model: config.openRouter.llmModel,
|
||||
temperature: 0,
|
||||
// Without this OpenRouter reserves the model's entire output window against
|
||||
// the key's remaining budget and rejects the call with a 402 before it ever
|
||||
// runs -- 131k tokens reserved to produce a few hundred. Reasoning models
|
||||
// also spend completion tokens on reasoning, so leave real headroom.
|
||||
max_tokens: Math.max(512, Number(config?.openRouter?.maxTokens || process.env.OPEN_ROUTER_MAX_TOKENS) || 8000),
|
||||
response_format: { type: 'json_object' },
|
||||
messages: [
|
||||
{ role: 'system', content: 'You are a coordinator. Extract only evidence-backed categorical hypotheses. Never output probabilities, expected returns, confidence scores, position sizes, or trade actions.' },
|
||||
|
||||
Reference in New Issue
Block a user