Benchmark Twitter is a poor CFO. DeepSeek V4 and Qwen-class models now sit near frontier quality on a lot of coding work at a fraction of Opus or GPT-5.6 Sol prices. If every internal tool hits a US flagship, you are not “quality-first”. You are un-routed.
A routing policy that survives finance
- Default: cheapest model that passes your evals.
- Escalate: only when a wrong answer is expensive.
- Self-host: when data cannot leave the building.
Pay Claude or GPT when the task is a nasty multi-file failure or a decision a client will remember. Pay DeepSeek or Qwen when you are generating tests, migrations, and first drafts. The arithmetic is in dollars per task, not dollars per million tokens.
This is how a studio that works Venezuela + USA stays competitive against a shop that burns $200/seat plus uncapped API. See the labor side in which engineering role vanished, and the delivery side in shipped work.