Ask Runable forDesign-Driven General AI AgentTry Runable For Free
Runable
Back to Blog
Technology6 min read

Nvidia wants to stop AI costs skyrocketing with its new software router — but will it really make a difference? | TechRadar

Routing to the ideal AI agent to save compute Discover insights about nvidia wants to stop ai costs skyrocketing with its new software router — but will it real

TechnologyInnovationBest PracticesGuideTutorial
Nvidia wants to stop AI costs skyrocketing with its new software router — but will it really make a difference? | TechRadar
Listen to Article
0:00
0:00
0:00

Nvidia wants to stop AI costs skyrocketing with its new software router — but will it really make a difference? | Tech Radar

Overview

News, deals, reviews, guides and more on the newest computing gadgets

Start exploring exclusive deals, expert advice and more

Details

Unlock and manage exclusive Techradar member rewards.

Unlock instant access to exclusive member features.

Get full access to premium articles, exclusive features and a growing list of member rewards.

Nvidia wants to stop AI costs skyrocketing with its new software router — but will it really make a difference?

When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.

Nvidia's open source Ne Mo Switchyard 'smartly' routes each agent request to the cheapest model that can handle it

The approach allows it to claim a 74% cost cut against a frontier-only model in return for a 6% reduction in accuracy

Ne Mo Switchyard joins an increasingly well-established territory that already sees players such as Route LLM, Lite LLM, and Open Router, in addition to in-house efforts at Open AI and AT&T to save costs

Nvidia has released Ne Mo Switchyard, an open source model router that sits between an application and a pool of language models and decides, per request or per turn, which model should handle a particular task.

The aim is to maximize efficiency by picking the right model for the right task: pushing frontier-level models to do simple tasks that much smaller or cheaper models could handle is not only inefficient, but also costly for enterprise customers.

This comes at a time when enterprise customers are already grappling with rising costs that can quickly spiral out of control as they increasingly adopt AI.

An open-source addition to a growing chorus of AI routers

The problem Nvidia is trying to solve isn't new, and industry giants are already attracting a lot of attention. The emerging field itself is potentially lucrative, even if it focuses on cutting costs, and major players are looking to get in on the action; Stripe's upcoming $7+ billion acquisition of Open Router, underscores this.

Switchyard is no different from its peers; it is essentially a proxy that uses a routing algorithm to decide which AI model to send queries to, while delivering an answer in the format that the calling application or user requires.

Why 95% of enterprise GPUs sit idle while AI startups can't get compute

Why Nvidia's Nemo Claw signals the true enterprise agent era

It accepts Open AI, Anthropic, and Responses API requests, translates between them, and documents the selected model, decision rationale, token usage, and latency for each call, enabling transparency and letting it tweak or tune its approach over time.

This makes sense compared with a one-size-fits-all approach that would otherwise be prevalent in a world with limited frontier-level AI compute, as LLMs continue to grow larger and more demanding. One example is what Nvidia points to: Nemotron Parse, a one-billion-parameter model primarily designed to extract structure from PDFs.

The question remains, however, whether such an approach is always feasible. Leveraging such a routing tool comes at a cost; Lang Chain's test of Nemotron 3.5 Lightning found that its judge model, which determines whether the agent is still on track after each turn, consumed up to 21.2% of total cost, second only to what it spent on Claude's Opus 4.8.

Running a more optimized judge model could therefore yield greater savings or lower overhead, but it could also hinder cases where users need a more basic AI model, and the judge model increases costs by reading or checking output at every turn.

A secondary concern would be variance: while costs are noticeably lower than running just Opus 4.8 in testing, it does swing from anywhere between

2.16and2.16 and
3.61, a mammoth 67% movement essentially guided by when it escalated queries to a bigger or smarter model. This tradeoff makes cost hard to predict: while routing lowers average spend, it widens the distribution around it.

‘Those two jobs need different physics’: Rebellions CEO says training and inference need different chips

Nvidia lays out its thoughts on how storage has become the next frontier of AI

Lang Chain also says users should not use Switchyard if their workloads are short or latency-sensitive; it estimates it adds about 700ms of latency because the judge model has to read all output.

Nvidia's take here is simple: it is pushing a packaged, open-source solution that offers a broad set of integrations in what is essentially a relatively fragmented space. Routing work to smaller open-weight models pushes inference toward hardware enterprises own, which Nvidia sells.

At the same time, cheaper inference per task has historically meant more inference, not less, which is the outcome a company selling accelerators would choose to woo enterprises, many of which are still on the fence when it comes to directly investing in the hardware that powers it all.

Follow Tech Radar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.

Having built hundreds of gaming PCs and being an avid gamer in his spare time, Rahim tends to have stronger opinions about hardware than most. This is particularly on display when he gets his way with powerful, but minimalistic RGB builds even as Small Form Factor (SFF) PCs come a close second.

You must confirm your public display name before commenting

Ubiquiti sued by Ukrainian families over claims its tech powered Russian battlefield drones

Ukraine reveals new UAV equipped with 24 hours of flight time and a range of 2,000 kilometres — enough to reach Siberia and back

'When I heard the released demo, I was shocked, angered and in disbelief' — Quote of the day by Scarlett Johansson on GPT-4's infamous 'Sky' launch voice

The US Marines have a new tactic to tackle drones — just shoot them down

'The attacks we found only scratch the surface of what is possible': Experts say so-called 'Proactive SIM' cards can hijack smartphones, Io T devices and even EV chargers

Tech Radar is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site.

© Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036.

Key Takeaways

  • News, deals, reviews, guides and more on the newest computing gadgets
  • Start exploring exclusive deals, expert advice and more
  • Unlock and manage exclusive Techradar member rewards
  • Unlock instant access to exclusive member features
  • Get full access to premium articles, exclusive features and a growing list of member rewards

Cut Costs with Runable

Cost savings are based on average monthly price per user for each app.

Which apps do you use?

Apps to replace

ChatGPTChatGPT
$20 / month
LovableLovable
$25 / month
Gamma AIGamma AI
$25 / month
HiggsFieldHiggsField
$49 / month
Leonardo AILeonardo AI
$12 / month
TOTAL$131 / month

Runable price = $9 / month

Saves $122 / month

Runable can save upto $1464 per year compared to the non-enterprise price of your apps.