The env var has been set to google/gemma-4-31b-it all along and nothing ever
mapped it onto openRouter.cheapModel, so graphWorker's fallback chain silently
used the main model instead. Graph entity resolution is the highest volume llm
call in the system and its entire job is to reply with the number of a match.
Measured on the live key, same prompt:
deepseek v4 flash 7564 completion tokens, 7558 of them reasoning $0.0013708
gemma-4-31b-it 14 completion tokens, 0 reasoning $0.0000121
113x. Gemma is actually the more expensive model per token, which is why this
was worth measuring rather than reasoning about prices: the cost is not the
price of the tokens, it is a reasoning model spending seven thousand tokens
thinking about a multiple choice question.
This also explains the reasoning tokens dominating the usage dashboard, and why
the daily spend roughly doubled today rather than yesterday. Removing the token
ceiling let a trivial prompt reason without bound. The ceiling was never the
right control for that, the model choice is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WnNxwxfXSbeNtjvtz5gayb