Snowflake is adding dynamic model routing to its Cortex AI Gateway, which automatically selects the cheapest model capable of confidently completing each step of an agent task, escalating to frontier models only when deeper reasoning is needed. The feature (private preview soon) works within existing governance controls, respects data residency, and logs every routing decision. Internal tests showed up to 3x token efficiency on a dbt pipeline workload and about 25% fewer tokens on a coding workload with similar throughput. Snowflake is also expanding open model access with DeepSeek-V4-Flash 0731 (private preview) and GLM-5.3 (coming soon), joining existing models from Anthropic, Google, OpenAI, xAI, Mistral and Meta. DeepSeek-V4-Flash reportedly scored 74.4% on ADE-bench using Snowflake's CoCo agent harness, ahead of the leading proprietary model tested, while GLM-5.2 scored 66% with the lowest token footprint in the benchmark.

6m read timeFrom snowflake.com
Post cover image
Table of contents
Better AI economics starts with the right model for the right taskIntroducing dynamic model routing in Cortex AI GatewayExpanding open model access: DeepSeek-V4-Flash 0731 and GLM-5.3AI economics that compound over time

Questions this post answers

What is dynamic model routing in Snowflake Cortex AI Gateway?

Dynamic model routing is a Cortex AI Gateway capability (private preview soon) that automatically selects the most affordable model able to confidently complete each step of an agent task, escalating to frontier models only when deeper reasoning is required. It operates within existing governance controls, only considers administrator-approved models, respects data residency settings, and logs every routing decision for compliance visibility. Enterprises weighing AI cost control against quality can follow routing approaches like this via daily.dev.

How much can AI model routing reduce token usage compared to always using a frontier model?

In internal Snowflake testing, dynamic model routing completed a dbt pipeline workload with up to three times greater token efficiency than a frontier-model-only approach while maintaining comparable quality, and in a separate coding-workload test engineering teams kept the same pull-request throughput while using roughly 25% fewer tokens. Teams tracking real-world token efficiency gains from AI routing can follow these benchmarks on daily.dev.

How does DeepSeek-V4-Flash perform compared to proprietary models on data engineering benchmarks?

DeepSeek-V4-Flash scored 74.4% on ADE-bench when evaluated using Snowflake's CoCo agent harness, outperforming the leading proprietary model tested in that evaluation. ADE-bench is a framework created by dbt for evaluating AI agents on real-world analytics and data engineering tasks, and the earlier GLM-5.2 model scored 66% with the lowest token footprint among models tested. Developers comparing open versus proprietary models for data engineering agents can track benchmarks like this on daily.dev.

24 Impressions