Snowflake targets AI costs with dynamic model routing

Snowflake is expanding its AI infrastructure proposition with dynamic model routing, designed to help enterprises control the economics of AI by automatically matching workloads with models based on quality, speed, customer preferences and cost.

The capability is being introduced through Cortex AI Gateway and Snowflake’s AI products, including Snowflake CoCo and Snowflake CoWork. It can also be used by third-party AI agents connected through Cortex AI Gateway.

The move comes as enterprises deploy a growing mix of AI applications and agents. Using a frontier model for every task can increase inference costs, while manually selecting and managing models for different workloads adds operational complexity.

Dynamic model routing is intended to address both problems. Lower-complexity or repetitive tasks can be directed towards more efficient models, while workloads requiring deeper reasoning can be routed to frontier models. Snowflake said this allows customers to reduce unnecessary inference spending without having to determine the appropriate model for every request.

“Enterprises are becoming much more rigorous about the economics of AI. The question is no longer how much AI they are using, but whether that AI is translating into meaningful business value,” said Sridhar Ramaswamy, CEO, Snowflake. “Achieving intelligence efficiency requires the flexibility to use the best model for each task as the landscape evolves.”

Snowflake is also expanding the model options available through Cortex AI with DeepSeek-V4-Flash 0731 and GLM-5.3. These will add to a portfolio that includes models from Anthropic, OpenAI, Google, SpaceXAI, Meta and Mistral.

The broader model selection is intended to give enterprises more flexibility in balancing performance and cost while retaining governance over enterprise data within Snowflake. The company said its approach is designed to accommodate the rapid evolution of open models without requiring customers to repeatedly rebuild applications or infrastructure.

Snowflake’s internal testing points to potential efficiency gains from using a mix of models rather than relying exclusively on frontier models. In one evaluation, agents using dynamic model routing to build a dbt pipeline achieved up to 3x greater token efficiency than a frontier-model-only approach while maintaining the same quality. In another test, engineering teams completed the same number of pull requests with 25% greater token efficiency.

The economics of AI are also being addressed through greater visibility and controls over consumption. Cortex AI Gateway provides administrators with visibility into token usage and costs, while Snowflake CoCo can extend controls through role-based access and tagging. Administrators can set default models, attribute consumption to teams or cost centres, establish user quotas and receive notifications as usage approaches defined limits.

Snowflake is positioning these capabilities around what it calls “intelligence efficiency”, a measure of how effectively organisations convert compute, models, data and context into business impact.

The company is also adding open models as their capabilities mature. Snowflake said its AI Research Team evaluated DeepSeek-V4-Flash on enterprise-focused tasks, with the model scoring 74.4% on data engineering tasks in its testing. GLM-5.2 scored 62.8% while using fewer tokens than the other models evaluated.

The shift towards model routing reflects a broader change in enterprise AI economics. As organisations move from experimentation to production, the question is increasingly not simply which model performs best, but whether the most capable model is necessary for every task.

For Snowflake, the strategy is to absorb that model-selection complexity at the platform layer while giving enterprises greater control over spending, model choice and governance as AI workloads scale.

AI
Comments (0)
Add Comment