Nvidia wants to stop AI costs skyrocketing with its new software router — but will it really make a difference? | Tech Radar
Overview
News, deals, reviews, guides and more on the newest computing gadgets
Start exploring exclusive deals, expert advice and more
Details
Unlock and manage exclusive Techradar member rewards.
Unlock instant access to exclusive member features.
Get full access to premium articles, exclusive features and a growing list of member rewards.
Nvidia wants to stop AI costs skyrocketing with its new software router — but will it really make a difference?
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
Nvidia's open source Ne Mo Switchyard 'smartly' routes each agent request to the cheapest model that can handle it
The approach allows it to claim a 74% cost cut against a frontier-only model in return for a 6% reduction in accuracy
Ne Mo Switchyard joins an increasingly well-established territory that already sees players such as Route LLM, Lite LLM, and Open Router, in addition to in-house efforts at Open AI and AT&T to save costs
Nvidia has released Ne Mo Switchyard, an open source model router that sits between an application and a pool of language models and decides, per request or per turn, which model should handle a particular task.
The aim is to maximize efficiency by picking the right model for the right task: pushing frontier-level models to do simple tasks that much smaller or cheaper models could handle is not only inefficient, but also costly for enterprise customers.
This comes at a time when enterprise customers are already grappling with rising costs that can quickly spiral out of control as they increasingly adopt AI.
An open-source addition to a growing chorus of AI routers
The problem Nvidia is trying to solve isn't new, and industry giants are already attracting a lot of attention. The emerging field itself is potentially lucrative, even if it focuses on cutting costs, and major players are looking to get in on the action; Stripe's upcoming $7+ billion acquisition of Open Router, underscores this.
Switchyard is no different from its peers; it is essentially a proxy that uses a routing algorithm to decide which AI model to send queries to, while delivering an answer in the format that the calling application or user requires.
Why 95% of enterprise GPUs sit idle while AI startups can't get compute
Why Nvidia's Nemo Claw signals the true enterprise agent era
It accepts Open AI, Anthropic, and Responses API requests, translates between them, and documents the selected model, decision rationale, token usage, and latency for each call, enabling transparency and letting it tweak or tune its approach over time.
This makes sense compared with a one-size-fits-all approach that would otherwise be prevalent in a world with limited frontier-level AI compute, as LLMs continue to grow larger and more demanding. One example is what Nvidia points to: Nemotron Parse, a one-billion-parameter model primarily designed to extract structure from PDFs.
The question remains, however, whether such an approach is always feasible. Leveraging such a routing tool comes at a cost; Lang Chain's test of Nemotron 3.5 Lightning found that its judge model, which determines whether the agent is still on track after each turn, consumed up to 21.2% of total cost, second only to what it spent on Claude's Opus 4.8.
Running a more optimized judge model could therefore yield greater savings or lower overhead, but it could also hinder cases where users need a more basic AI model, and the judge model increases costs by reading or checking output at every turn.
A secondary concern would be variance: while costs are noticeably lower than running just Opus 4.8 in testing, it does swing from anywhere between
‘Those two jobs need different physics’: Rebellions CEO says training and inference need different chips
Nvidia lays out its thoughts on how storage has become the next frontier of AI
Lang Chain also says users should not use Switchyard if their workloads are short or latency-sensitive; it estimates it adds about 700ms of latency because the judge model has to read all output.
Nvidia's take here is simple: it is pushing a packaged, open-source solution that offers a broad set of integrations in what is essentially a relatively fragmented space. Routing work to smaller open-weight models pushes inference toward hardware enterprises own, which Nvidia sells.
At the same time, cheaper inference per task has historically meant more inference, not less, which is the outcome a company selling accelerators would choose to woo enterprises, many of which are still on the fence when it comes to directly investing in the hardware that powers it all.
Follow Tech Radar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
Having built hundreds of gaming PCs and being an avid gamer in his spare time, Rahim tends to have stronger opinions about hardware than most. This is particularly on display when he gets his way with powerful, but minimalistic RGB builds even as Small Form Factor (SFF) PCs come a close second.
You must confirm your public display name before commenting
Ubiquiti sued by Ukrainian families over claims its tech powered Russian battlefield drones
Ukraine reveals new UAV equipped with 24 hours of flight time and a range of 2,000 kilometres — enough to reach Siberia and back
'When I heard the released demo, I was shocked, angered and in disbelief' — Quote of the day by Scarlett Johansson on GPT-4's infamous 'Sky' launch voice
The US Marines have a new tactic to tackle drones — just shoot them down
'The attacks we found only scratch the surface of what is possible': Experts say so-called 'Proactive SIM' cards can hijack smartphones, Io T devices and even EV chargers
Tech Radar is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site.
© Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036.
Key Takeaways
- News, deals, reviews, guides and more on the newest computing gadgets
- Start exploring exclusive deals, expert advice and more
- Unlock and manage exclusive Techradar member rewards
- Unlock instant access to exclusive member features
- Get full access to premium articles, exclusive features and a growing list of member rewards



