惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
DataBreaches.Net
罗磊的独立博客
雷峰网
雷峰网
量子位
V
Visual Studio Blog
Vercel News
Vercel News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Cloudflare Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
宝玉的分享
宝玉的分享
月光博客
月光博客
Martin Fowler
Martin Fowler
aimingoo的专栏
aimingoo的专栏
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Microsoft Security Blog
Microsoft Security Blog
博客园 - 叶小钗
腾讯CDC
Engineering at Meta
Engineering at Meta
博客园 - Franky
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Y
Y Combinator Blog
Recent Announcements
Recent Announcements
Jina AI
Jina AI
A
About on SuperTechFans

Latest from TechRadar in Pro

VodafoneThree gets Ofcom approval to bring satellite connectivity to your smartphone Is this the tipping point for AI at work? New Gallup survey finds half of all US employees now use it in some way 'Every Apple user needs to know about this nasty scam': Fake warnings tell users their iCloud data will be… 'Makes it even more disappointing': Microsoft backs fossil fuel big time with $7 billion deal in race for AI… 'Maybe it’s not science fiction': Solar panels are causing rainwater to fall in one of the driest places… Maine becomes first US state to pass data centre construction ban Dozens of WordPress plugins hijacked to target thousands of sites Drone-killing laser weapons greenlit for use in US airspace – FAA and Defense Department say high-energy weapons are ‘ready to protect all air travelers from illicit drone use’ despite airspace restrictions and friendly-fire incidents 'We are currently being extorted' — crypto giant Kraken says it is facing extortion attack, here's… I tried 7 free MTD software – now I've ranked my top picks as a freelancer Jackery McGraw Hill becomes latest to see its Salesforce data hacked Looking for a new PC? Now might be great time to upgrade, as Gartner figures claim shipments are rising — while… The new engineering playbook: how AI design copilots are reshaping product development Farewell Surface Hub — Microsoft kills off its super-sized touchscreen displays, but you might still be able to get one if you act fast 'We have no interest in patient data in the UK': Palantir UK head defends record as criticisms rise Amazon’s new AI Bio Discovery tool can provide ‘every researcher’ with ‘lab-in-the-loop drug discovery’ – 40+ AI biology models can filter 300,000 novel antibody candidates down to the top results for testing in just weeks Over 100 Chrome Web Store extensions found stealing user data from thousands of accounts Europe wants tech sovereignty but is this realistic? Enterprise AI governance cannot live in a prompt. So where is the safety net? Why 2026 is the year of flexibility without friction: solving the multi-platform crisis OpenAI reveals its Mythos rival designed for cybersecurity pros When cyberattacks are inevitable, recovery becomes the strategy Closing the cloud complexity gap LaLiga uses AI to fight illegal streaming that costs its clubs $800m a year Intel and Google expand long-term chip partnership to power AI systems 'Chatbots respond not just to what you ask, but how you ask it': Report finds AI agents might be sucking up to… 'Smartphones have physical limitations': Report explains why AI is kickstarting a billion-dollar hardware arms… 'I’m pretty sure actually we really do not need to work for five days' Zoom CEO calls for end of traditional work schedules — says 3-day working week should become the norm 'It's more common than you think': Experts reveal how hackers are trying to hijack your inbox with these…
Why businesses are shifting from cloud to on-prem amid th...
Michael Jin · 2026-04-30 · via Latest from TechRadar in Pro

Offering speed, flexibility, and the ability to scale without heavy upfront investment, the public cloud has for years been the model of efficiency. But as AI becomes embedded across every function of organizations, what once seemed like convenience now looks a lot more like a permanent cost burden.

That’s why many businesses are shifting from a cloud-first mindset toward a more balanced, hybrid approach, one that sees AI workloads brought back on-premise

Senior Product Director of MINISFORUM.

Cloud used to be a major cost saver, but in 2026, the economics are changing quickly. Ingress and egress fees, combined with the premium charged for GPU compute cycles, have ballooned as more AI models run.

Article continues below

When 10% of top-line revenue goes to a cloud provider just to keep the lights on, organizations feel like they’re not simply renting infrastructure but paying a recurring tax on their own growth.

This is the state of play with the always-on nature of today’s AI models.

Frequent, high-volume tasks are driving cost increases. Enterprises are now using large language models (LLMs) to summarize internal meetings, scan customer support tickets, and run continuous retrieval-augmented generation (RAG) pipelines.

Individually, these API calls seem inexpensive. But at scale, they are a massive recurring expense. AI agents bring more complexity. These systems function more like digital employees, planning tasks, verifying outputs and retrying workflows.

Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!

From renting to owning

With public cloud pricing models, the more a team relies on AI, the more an organization pays. In other words, there’s a tax on realizing AI’s full potential.

On-prem infrastructure turns that upside-down. A one-time investment in high-performance hardware converts unpredictable monthly expenses into fixed, depreciable assets. Companies own the computing capability outright rather than paying exorbitant rent.

The cost of local hardware is often recouped quickly when compared to ongoing API usage or GPU rental fees, particularly for predictable, always-on workloads.

But cost is just part of the equation. Performance is the other.

In the cloud, workloads typically run on shared infrastructure. Organizations often operate on a “slice” of a server alongside other tenants, introducing latency, resource contention, and performance variability.

By contrast, local AI runs on dedicated hardware. There is no network lag, no shared queues, and no “noisy neighbor” interference. For end users, that translates into immediate responsiveness.

The governance imperative

Data sovereignty is another driver of the on-prem trend.

In a public cloud environment, sensitive data resides on third-party infrastructure, creating challenges for compliance, auditing, and intellectual property protection.

On-prem AI changes that dynamic. Prompts, proprietary training data, and outputs remain within the organization’s physical and logical boundaries. Compliance with frameworks like GDPR or HIPAA becomes more straightforward because data residency is guaranteed by design.

This also addresses growing concerns around “prompt leaks.” When employees input sensitive information into external AI systems, there is a risk of unintended persistence or exposure. Localized AI environments create a controlled, secure environment for experimentation and deployment.

Smaller, more efficient models are making this possible. Businesses do not need hyperscale infrastructure for every use case.

That’s why we are beginning to see the “rightsizing” of AI. Capable assistants can now run on systems with 64GB or 128GB of high-speed memory. What once required a large, expensive server can now be done with a compact, cost-effective workstation.

Hybrid model

This transition to on-prem AI does not mean abandoning the cloud.

For most forward-looking businesses, the right solution is a hybrid model. Cloud can be used more strategically, reserved for large-scale training jobs and burst workloads that require massive, synchronized GPU resources.

At the same time, local infrastructure handles agentic AI programs, internal copilots, and sensitive data analysis.

As a strategic hub rather than a peripheral, companies can build environments that are faster, more secure, and more cost-efficient than a cloud-only approach.

They can attain full control over their data, eliminate hidden costs such as egress fees, and offer their teams a better experience.

In the future, we will see one person directing a team of agents, and in an enterprise, hundreds or even thousands of agents may continuously plan, call tools, share context, verify results, and retry tasks — all of which drive token usage sharply higher. This is a fundamental shift in how AI is used.

Collectively, these trends point to the emergence of a “private AI” model.

The shift from cloud-first to hybrid and on-prem AI is being driven by a convergence of forces: economics, governance, and performance. In 2026, the question is no longer whether to use the cloud, but how to use it strategically while keeping control over the workflows that matter most.

We've featured the best cloud computing provider.

This article was produced as part of TechRadar Pro Perspectives, our channel to feature the best and brightest minds in the technology industry today.

The views expressed here are those of the author and are not necessarily those of TechRadarPro or Future plc. If you are interested in contributing find out more here: https://www.techradar.com/pro/perspectives-how-to-submit

Senior Product Director of MINISFORUM.