Ask Runable forDesign-Driven General AI AgentTry Runable For Free
Runable
Back to Blog
Technology7 min read

Beware the token trap: Why saving on inference might put your ADLC at risk | TechRadar

Token use can create unexpected, sizeable costs for organizations Discover insights about beware the token trap: why saving on inference might put your adlc at

TechnologyInnovationBest PracticesGuideTutorial
Beware the token trap: Why saving on inference might put your ADLC at risk | TechRadar
Listen to Article
0:00
0:00
0:00

Beware the token trap: Why saving on inference might put your ADLC at risk | Tech Radar

Overview

News, deals, reviews, guides and more on the newest computing gadgets

Start exploring exclusive deals, expert advice and more

Details

Unlock and manage exclusive Techradar member rewards.

Unlock instant access to exclusive member features.

Get full access to premium articles, exclusive features and a growing list of member rewards.

Beware the token trap: Why saving on inference might put your ADLC at risk

Token use can create unexpected, sizeable costs for organizations

When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.

Agentic AI’s prolific use of tokens can create sizeable, unexpected costs for organizations. But saving on token costs without factoring in risk can be a fatal step.

As upfront prices for flagship artificial intelligence models continue to shrink, organizations have begun to wise up to the hidden costs they encounter with agentic AI models.

Specifically, the costs of tokens, which may look tiny when viewed as individual charges, can add up exponentially as AI agents become more active, leaving organizations with hefty AI expenditures they may not have anticipated.

This is putting CISOs in something of a bind. If they seek to save money on inference costs, primarily driven by token generation incurred by agentic AI, they may increase their security risk and accumulate hidden technical debt that puts their Agentic Development Lifecycle (ADLC) in jeopardy. It’s a problem that many CISOs may not have factored into their security budgets, but it cannot be left unaddressed.

The effectiveness of automated security processes is being impeded by fragmented pricing across the AI landscape, whether we’re talking about hyper-optimized nano models (essentially lightweight, yet powerful models built for a specific use, like Google’s Nano Banana 2 image generator) or premium reasoning engines, like Salesforce Atlas or Open AI o 3.

Beyond Tokenmaxxing: the rising token tax on enterprise AI

Tokenomics: AI Has an income statement - it’s time to read it

What is Tokenmaxxing, and why should businesses care about it?

Organizations do have to keep a close eye on token costs to prevent them from spiraling, but CISOs also need to examine how agentic AI is affecting their security.

Erratic pricing has been a trademark of generative AI pretty much from the beginning.

About two years after Open AI released Chat GPT, the Chinese company Deep Seek shook up the AI market with the release of a powerful, open-weighted large language model whose training parameters were publicly available, allowing users to customize the model and build on the cheap compared with other generative AI models.

Chat GPT-maker Open AI and other AI companies started doing the same, and suddenly, the costs of using Gen AI systems dropped off a cliff. In fact, prices fell faster for Gen AI than for any other technology in history.

The emergence of agentic AI has introduced some stealth costs into the equation, however. The costs of agentic software range from free for open-source models to enterprise agents, with prices that vary from one-time fees (roughly

15,000forbasicmodelstomorethan15,000 for basic models to more than
1 million for global enterprise models) to monthly subscriptions (which can range from a few thousand to $13,000 or more).

Token maxxing is your AI program’s quiet failure mode

How to embrace the spirit of ‘Tokenmaxxing’ without breaking the bank

But those costs are fixed. Inference costs are another story: they scale with usage and can amount to 90% of AI lifecycle costs.

Tokens come into play when an AI agent requests processing from Gen AI models, which charge agents for processing information. At a glance, the costs may appear inconsequential. Input tokens generally range from 15 cents to

5permillionrequests.Outputtokens,whichrequireslightlymoreprocessing,costfromabout60centsto5 per million requests. Output tokens, which require slightly more processing, cost from about 60 cents to
25 per million.

They may start small, but can add up in no time, thanks to AI agents that work very quickly, autonomously, and unpredictably. They are designed to interact with systems and other agents throughout the enterprise. A single action might generate scores of LLM calls. Token use, which has grown exponentially with the use of AI agents, has already increased IT budgets by about 20% according to recent estimates.

The accelerating cost of agentic AI is prompting CISOs to look for ways to save money where they can, and one way is to identify LLMs that charge the least per token. But what they may not be considering are the risk factors associated with those LLMs. If CISOs concern themselves only with the costs, they may open themselves up to security risks.

But better security doesn’t necessarily have to cost more. Depending on what they’re using agentic AI for, they may find they don’t always have to trade security for lower token costs.

There are a few things organizations can do to help stop token costs from getting out of hand, including:

Match Agents and LLMs to the Job at Hand. Commodity AI systems can cost little or nothing, but they lack the deep reasoning for complex security synthesis. But not every application or function within the organization requires a reasoning engine. You can set up agents to work with low-cost LLMs on low-risk projects, while preserving higher-cost LLMs for critical tasks. It’s also worth being aware of which agents are likely to request more LLM calls.

Factor Risk Scores in Choosing Agents and LLMs. The security implications of using AI can’t be ignored. When developing a budget plan, include risk factors.

Monitor Workflows. Keeping a close watch on workflows can help you track costs and performance, allowing you to better understand which tools work best in which situations.

Lean on Human Oversight. Despite agentic AI’s autonomy, in fact, because of agentic AI’s autonomy, forgetting about the importance of the human element is risky business. Teams need thorough upskilling in secure development, with clearly defined ownership roles. And they must be given prominent oversight roles throughout the ADLC.

Agentic AI is fast becoming integral to enterprise operations, and organizations must control its associated costs. But a race to the bottom on token pricing creates hidden technical debt. Instead, CISOs need to weigh security performance when choosing AI tools as part of establishing an up-to-date security maturity model and an AI governance policy that emphasizes performance, costs, and risk management.

Only that approach allows agentic AI to be deployed without either breaking the budget or putting your entire organization at risk.

This article was produced as part of Tech Radar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today.

The views expressed here are those of the author and are not necessarily those of Tech Radar Pro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit

Pieter Danhieux, Co-Founder and CEO, Secure Code Warrior.

You must confirm your public display name before commenting

'We failed you today': Namecheap down for several hours after a data center cooling failure, leaving customers furious

Stuart Fails to Save the Universe Big Bang Theory cameos based on who was available

Experts warn Chat GPT isn't just predicting words anymore — It's now predicting human thoughts

Quordle hints and answers for Friday, August 14 (game #1663)

NYT Connections hints and answers for Friday, August 14 (game #1160)

Tech Radar is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site.

© Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036.

Key Takeaways

  • News, deals, reviews, guides and more on the newest computing gadgets
  • Start exploring exclusive deals, expert advice and more
  • Unlock and manage exclusive Techradar member rewards
  • Unlock instant access to exclusive member features
  • Get full access to premium articles, exclusive features and a growing list of member rewards

Cut Costs with Runable

Cost savings are based on average monthly price per user for each app.

Which apps do you use?

Apps to replace

ChatGPTChatGPT
$20 / month
LovableLovable
$25 / month
Gamma AIGamma AI
$25 / month
HiggsFieldHiggsField
$49 / month
Leonardo AILeonardo AI
$12 / month
TOTAL$131 / month

Runable price = $9 / month

Saves $122 / month

Runable can save upto $1464 per year compared to the non-enterprise price of your apps.