Coinbase AI Spend: Switching to GLM and Kimi
For anyone optimizing an AI workflow, this is a huge signal. It shows that diversifying your LLM agent strategy—rather than sticking to a single expensive ecosystem—is the most effective way to scale without burning through your budget. If a major fintech player is comfortable migrating critical infrastructure to these models to slash overhead, it's time for the rest of us to stop overlooking high-efficiency alternatives.
This is a practical lesson in deployment: don't overpay for brand names if a more efficient model handles the specific task just as well. Moving to a multi-model architecture allows you to route simple queries to cheaper models while reserving the "heavy hitters" for complex reasoning, which is likely how they achieved such a steep drop in spending.