Ask Runable forDesign-Driven General AI AgentTry Runable For Free
Runable
Back to Blog
Technology6 min read

Compression’s new goal: Reducing how much an AI ‘overthinks’ | TechRadar

Compression is now a pillar of operational AI Discover insights about compression’s new goal: reducing how much an ai ‘overthinks’ | techradar........

TechnologyInnovationBest PracticesGuideTutorial
Compression’s new goal: Reducing how much an AI ‘overthinks’ | TechRadar
Listen to Article
0:00
0:00
0:00

Compression’s new goal: Reducing how much an AI ‘overthinks’ | Tech Radar

Overview

News, deals, reviews, guides and more on the newest computing gadgets

Start exploring exclusive deals, expert advice and more

Details

Unlock and manage exclusive Techradar member rewards.

Unlock instant access to exclusive member features.

Get full access to premium articles, exclusive features and a growing list of member rewards.

Compression’s new goal: Reducing how much an AI ‘overthinks’

When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.

Back in the late ‘90s, you compressed because storage was limited, bandwidth was expensive, and users valued rapid response.

Then, file compression was about encoding, restructuring or modifying data to reduce its size – smaller payloads meant faster, more efficient delivery and less storage space.

Traditionally, compression was about performance. Then it was about bandwidth. But the AI era has flipped our long-standing assumptions of compression on its head.

Google’s new compression drastically shrinks AI memory use while quietly speeding up performance

“Rewriting the blueprint, not removing bricks”: Multiverse Computing says it can shrink large AI models and cut memory use in half

Why AI must shrink to reach its enterprise potential

Distinguished Engineer in the Office of the CTO at F5.

Today, compression is about not bankrupting yourself on inference.

In the AI world, every token generated is an act of cognition and cognition, for machines, is expensive. So, we no longer compress to make things smaller. We compress so it is cheaper for AI to “think.”

And yes, bandwidth still costs money. Cloud provider egress is infamous, and data transfer bills can still produce heart palpitations. But be honest and compare the cost of moving a megabyte across the wire with the cost of generating 10,000 tokens on a top-shelf large language model (LLM).

One is a forgotten rounding error on the monthly bill. The other is a sternly worded message from finance asking why you’ve suddenly consumed the budget for Q3.

Compression has flipped from optimization to cost control

It used to be that you optimized network paths, minimized payloads, and pre-compressed assets so your application wouldn’t take six days to load on a 3G connection. But LLMs have redefined bottlenecks in ways that feel almost disrespectful to the past three decades of systems engineering. Now the slowest, most expensive component in the system isn’t the network at all. It’s the brain.

The cost of generating text now dwarfs the cost of transporting it. Every token an LLM emits demands GPU cycles, VRAM, energy and latency. This isn’t cheap, and depending on your model of choice for the quarter, this is downright expensive. Because of this, the compression value chain has been inverted.

We now compress not to shrink the data, but to reduce the number of “thoughts” an AI has to “think”.

The post-transformer era has an answer to AI’s energy crisis

Why businesses are shifting from cloud to on-prem amid the agent boom

Compression used to live at the edge of the network in specialized devices. Then, it consolidated on application delivery controllers, taking on names like “minification” and “HTTP compression.” For a time, it was specialized functionality. Fast forward to today and it’s just part and parcel of application delivery.

But, thanks to AI tools, we’re seeing the emergence of new compression techniques. We’re no longer just compressing text using well-known algorithms. We’re striking out words like a Chicago- or AP-style editor with a pen full of red ink and something to prove.

Prompt compression has emerged as the new heavyweight champion. You shrink the prompt to shrink the invoice. Irrelevant details? Gone. Redundant context? Deleted. Overly chatty instructions? Trimmed like an overgrown hedge. The shorter the prompt, the fewer tokens consumed, and the happier your procurement department.

“Be concise” has quietly graduated from a writing preference to a cost-control strategy. Short answer = cheap answer. Long answer = someone’s paying for that verbosity. This is output compression.

Embedding compression is not about reducing bytes, it’s about reducing dimensionality. This reduces memory footprint, retrieval cost, and everything your vector store is quietly billing you for every minute.

Pruning, quantization and distillation are the foundations of model compression. In another era, these were academic curiosities. Today, they serve one purpose: to run it cheaper. If it also runs faster? Wonderful. If it fits on a smaller GPU? Miraculous. But the point is, and always has been, to lower the compute burn.

Compression is no longer a nicety; it’s a pillar of operational AI. Today, network is cheap. Storage is cheap. CPU is cheap. Memory is cheap enough that we barely pretend to manage it anymore. But GPU inference? That’s the new oil. And like oil, we now have a global economy dedicated to extracting every last drop efficiently.

It’s how you stay inside budget, scale responsibly, prevent accidental million-dollar token overruns, and prevent agents from rewriting War and Peace because you forgot to set max tokens. When your system’s most expensive operation is thinking, you start treating thoughts like a limited resource.

We compress now not because our networks can’t handle the load, but because our AIs can’t handle the invoice. Compression no longer serves the network. It serves the ledger. The future isn’t about making data smaller; it’s about making thinking cheaper.

This article was produced as part of Tech Radar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today.

The views expressed here are those of the author and are not necessarily those of Tech Radar Pro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit

Distinguished Engineer in the Office of the CTO at F5.

You must confirm your public display name before commenting

1VPN deal of the week: Get Norton VPN Standard at half-price with Tech Radar's exclusive deal where you get a $30 Amazon gift card with any two-year Norton VPN plan – perfect for keeping data secure on your phone, tablet, or computer

2 These bookshelf speakers just replaced basically every part of my hi-fi set-up in one fell swoop — after months of testing, I really appreciate the Edifier M90's sheer range of connectivity

3 Rivals season 2 review: Jilly Cooper bonkbuster is still as steamy as ever — but troubling tensions take over in first three-episode release

4 Quordle hints and answers for Monday, May 11 (game #1568)

5NYT Connections hints and answers for Monday, May 11 (game #1065)

Tech Radar is part of Future US Inc, an international media group and leading digital publisher. Visit our corporate site.

© Future US, Inc. Full 7th Floor, 130 West 42nd Street, New York, NY 10036.

Key Takeaways

  • News, deals, reviews, guides and more on the newest computing gadgets
  • Start exploring exclusive deals, expert advice and more
  • Unlock and manage exclusive Techradar member rewards
  • Unlock instant access to exclusive member features
  • Get full access to premium articles, exclusive features and a growing list of member rewards

Cut Costs with Runable

Cost savings are based on average monthly price per user for each app.

Which apps do you use?

Apps to replace

ChatGPTChatGPT
$20 / month
LovableLovable
$25 / month
Gamma AIGamma AI
$25 / month
HiggsFieldHiggsField
$49 / month
Leonardo AILeonardo AI
$12 / month
TOTAL$131 / month

Runable price = $9 / month

Saves $122 / month

Runable can save upto $1464 per year compared to the non-enterprise price of your apps.