Neither the coordinator nor the graph resolver set max_tokens, so openrouter
reserved the model's entire output window against the key's remaining budget and
returned 402 before running anything: 131k tokens reserved to produce a few
hundred. It never showed up with the old model because its output window is
small enough to fit under the limit.
Both are now bounded and overridable by env. The headroom is deliberate, the
newer reasoning models spend completion tokens thinking before they answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb