Connect with us

AI

Snowflake Uses Dynamic Model Routing to Cheapen Every Agent Task

Snowflake will route each Cortex task to a cheaper approved model, serve DeepSeek-V4-Flash next to warehouse data, and bill the work in AI Credits.

Published

on

Snowflake on Aug. 18, 2026 added dynamic model routing in Cortex AI Gateway so each agent step can use a cheaper model. The company will also put DeepSeek-V4-Flash 0731 and GLM-5.3 on the same governed path as Anthropic, OpenAI, Google, SpaceXAI, Meta, and Mistral. CEO Sridhar Ramaswamy’s own case is not a smaller bill for its own sake. As the cost of a task falls, more jobs become cheap enough to automate, which is how CoCo and CoWork keep spreading inside the data cloud.

Cortex AI Gateway Will Pick the Model for Each Step

Cortex AI Gateway, first shown on July 28, 2026, already sat between agents and the models, tools, and data they touch. Dynamic model routing plugs into that layer and into Snowflake CoCo, the coding agent, and Snowflake CoWork, the agent for business users. Third-party agents that already call the gateway get the same picker. The feature “will be private preview soon,” according to the launch note.

The gateway is told to take the most affordable approved model that can finish the step with confidence. Repetitive work can go to a smaller model. Deeper reasoning can still go to a frontier model. Admins, not app teams, decide which names are even in the pool, which matters for shops that have residency rules or a short list of allowed labs.

Ramaswamy wrote that customers set those approvals and the tradeoffs they care about, after which the gateway scores each task against policy plus live cost and performance data. After one model finishes, another model grades the result, and that loop is supposed to improve later picks as prices and quality move.

HOW THE GATEWAY CHOOSES

  • Approved list: Only models an admin has allowed are candidates for a request.
  • Residency: Existing data-residency settings still bind, so a pick cannot wander outside the geography the account already set.
  • Audit log: Each routing decision is recorded, including which model handled which request.
  • Quality loop: A second model scores the first model’s output and feeds that result back into later routing.
  • Spend caps: The gateway can show token use and cost, and CoCo can hang quotas, tags, and limit alerts on Snowflake’s usual roles.

That last piece is the CFO surface. CoCo can set a default model, push usage onto a team or cost center, and warn when a person is near a cap. The router then tries to make the remaining budget cover more finished work rather than more wasted frontier calls.

74.4% on dbt’s ADE-Bench for Data Work

A router is only as useful as the cheap models on the other end of it. Snowflake’s AI research team ran DeepSeek-V4-Flash on enterprise tasks using ADE-bench, a dbt framework for scoring agents on analytics and data-engineering work, with CoCo as the agent harness. DeepSeek-V4-Flash scored 74.4 percent and beat the leading paid models in that test. GLM-5.2, from an earlier run, scored 62.8 percent and used fewer tokens than any other model in the set. GLM-5.3 is the one Snowflake is adding; it was not the scored row.

SNOWFLAKE’S OWN ADE-BENCH SCORES

Model Status in Cortex What Snowflake published
DeepSeek-V4-Flash 0731 Private preview in CoCo 74.4% on ADE-bench data engineering, ahead of the paid models in the test
GLM-5.3 Private preview soon, may slip with model supply No ADE-bench figure in the launch note
GLM-5.2 Prior test only, not the new catalog row 62.8%, fewest tokens of any model in that test

Those scores are Snowflake’s, on Snowflake’s harness, on one bench. They are still the load-bearing claim hiding under the routing headline: a current open model already cleared the data-engineering bar the company uses for CoCo. Directories that DeepSeek V4 Flash and GLM 5.3 compared side by side are the public catalog Snowflake is trying to pull behind its own controls, not a second integration project for every app team.

Internal routing tests, which Snowflake says it will share on request and which can move with the workload, point the same way. Agents that used the mix built a dbt pipeline with up to 3x greater token efficiency than a frontier-only path, at comparable quality. In a separate coding test, teams shipped the same number of pull requests with 25 percent greater token efficiency.

A Router Only Helps If the Cheap Models Are Good

Model routing is not a Snowflake invention. Sanjeev Mohan, principal and founder of SanjMo, said the pain is the overhead of picking a model for every task at scale, and that a new lab release should not force a rebuild. He also treats the algorithm as table stakes. Microsoft, Databricks, OpenRouter, and Vercel already sell some form of the same idea. What Snowflake is selling is a picker that never leaves the controls already sitting on the warehouse.

Enterprises are drowning in model choices, but the real problem isn’t which model to pick. It’s the operational overhead of picking the right one for every task, at scale. Snowflake’s dynamic model routing directly addresses that gap. By automating intelligent model selection within Cortex AI Gateway, Snowflake is removing a real friction point that has been slowing enterprise AI deployment. The ability to match workload complexity to model cost, without rebuilding your infrastructure every time a new model drops, is exactly the kind of efficiency enterprises need to move from AI experimentation to AI at scale.

Sanjeev Mohan, Principal and Founder, SanjMo

That is why the DeepSeek row matters more than the marketing name “intelligence efficiency,” Ramaswamy’s phrase for turning compute, models, data, and context into a business result. If the cheap model cannot clear CoCo’s data work, the router has nothing safe to send the boring steps to, and every agent still burns frontier tokens.

Snowflake Runs the Open Weights Next to the Data

DeepSeek-V4-Flash 0731 is in private preview in CoCo now. GLM-5.3 is “coming soon,” self-hostable, and flagged as subject to change based on model supply. Product managers Siddharth Dwivedi, Paras Mehra, and Albert Cheng wrote that Snowflake serves these open models itself rather than proxying a third-party API, so inference stays inside the same perimeter as the tables.

They list three consequences. Inference sits next to governed data, which cuts transfer cost and extra latency. Snowflake can tune serving for the access patterns its own agents actually produce. Open models then inherit the same role-based access and audit trail that already wrap the warehouse, instead of a one-off exception for a lab’s public endpoint.

That is the lock customers should price, not the routing slogan. Leave later, and the picker, the log, and the in-house weights all have to be rebuilt somewhere else. Stay, and a new open model is a catalog row plus a routing policy, not a new vendor review for every CoCo skill.

What AI Credits Cost on Global Versus Regional Routing

Cheaper tokens only show up on the invoice if the meter says so. Snowflake bills CoCo, CoWork, Cortex Agents, and related AI features in AI Credits, a flat rate that does not change with edition. The rate still changes with where the request is allowed to run, and the per-token rate still depends on the model the orchestrator picked. Capacity discounts do not apply to AI Credits.

AI CREDIT PRICE BY ROUTING TYPE

Routing type Price per AI Credit When it applies
Global routing $2.00 Cross-region set to ANY_REGION, AWS_GLOBAL, GCP_GLOBAL, or AZURE_GLOBAL
Regional routing $2.20 Cross-region disabled, or pinned to a geography such as AWS_US or AZURE_EU

Snowflake’s own worked example is an Enterprise account that still pays $3.00 per Platform Credit. One hundred Platform Credits cost $300. One hundred AI Credits at $2.00 per AI Credit on global routing cost $200. Combined spend is $500. Regional routing raises the AI line to $2.20 a credit. CoWork and Cortex Agents add up across every tool the agent calls, so a sloppy multi-step job still stacks even after the router starts sending easy steps to DeepSeek.

On the Sept. 2, 2026 fiscal 2027 second-quarter call, Ramaswamy said CoCo had surpassed 9,100 accounts and CoWork 5,800. That is a lot of agent traffic to put on a $2.00 meter. A 25 percent token cut on pull requests, or up to 3x on a dbt pipeline, is how those accounts keep running without the quota emails arriving at noon.

Practitioners in Tokyo Are Still Waiting to Try It

The Aug. 18 note is a ship date for a policy, not a date every customer can click. Dynamic model routing is still “private preview soon.” GLM-5.3 is the same, and Snowflake warned that row can move. Only DeepSeek-V4-Flash 0731 is described as in private preview now, and only in CoCo.

THE PATH TO A GOVERNED ROUTER

  1. May 2026: Snowflake buys Natoma Labs, the MCP governance shop later folded into the gateway.
  2. July 28, 2026: Cortex AI Gateway is announced as the control layer for Snowflake agents and for third-party agents that plug in.
  3. Aug. 18, 2026: Dynamic model routing is added on paper, with DeepSeek-V4-Flash 0731 in private preview and GLM-5.3 still queued.
  4. Sept. 2, 2026: Q2 figures put CoCo above 9,100 accounts and CoWork at 5,800.
  5. Sept. 11, 2026: At the Tokyo world-tour stop, working Snowflake practitioners are still asking when they can actually use the gateway and CoCo support.

That Tokyo queue is the honest status. The people who already run warehouses treat Cortex AI Gateway as a thing they want in their hands, not a thing they have. Until the preview list opens, the 74.4 percent score and the up to 3x dbt result remain Snowflake’s lab, not a customer’s FinOps dashboard.

Cheaper Tasks Are How More Agents Get In

Ramaswamy’s launch line was that more AI use is not the same as more impact. He then described the product Snowflake actually wants to own: a flexible pool of models, a router that keeps rewriting the map as labs reprice, and a platform that hides that churn from the person in CoWork.

Enterprises are becoming much more rigorous about the economics of AI. The question is no longer how much AI they are using, but whether that AI is translating into meaningful business value. Achieving intelligence efficiency requires the flexibility to use the best model for each task as the landscape evolves. Snowflake’s role is to absorb that complexity so customers can focus on outcomes while we optimize model choice underneath.

Sridhar Ramaswamy, CEO, Snowflake

He also said the quiet part in the longer note. Open models already do a lot of simple, well-defined work at a fraction of frontier cost. When that cost per task falls, work that was too expensive to automate starts to clear the bar. Dynamic model routing is how Snowflake tries to make that true inside CoCo and CoWork without asking every team to hard-code DeepSeek this month and GLM the next. The bill still runs in AI Credits. The bet is that a cheaper step is a step a company will allow to run at all.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending