惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
CXSECURITY Database RSS Feed - CXSecurity.com
A
About on SuperTechFans
H
Help Net Security
Engineering at Meta
Engineering at Meta
G
Google Developers Blog
aimingoo的专栏
aimingoo的专栏
The Register - Security
The Register - Security
WordPress大学
WordPress大学
MongoDB | Blog
MongoDB | Blog
Hugging Face - Blog
Hugging Face - Blog
爱范儿
爱范儿
C
Check Point Blog
人人都是产品经理
人人都是产品经理
云风的 BLOG
云风的 BLOG
N
Netflix TechBlog - Medium
酷 壳 – CoolShell
酷 壳 – CoolShell
雷峰网
雷峰网
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - Franky
I
InfoQ
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
The GitHub Blog
The GitHub Blog
Last Week in AI
Last Week in AI
D
Docker
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 叶小钗
Jina AI
Jina AI
F
Fortinet All Blogs
宝玉的分享
宝玉的分享
小众软件
小众软件
有赞技术团队
有赞技术团队
F
Full Disclosure
月光博客
月光博客
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Recorded Future
Recorded Future
Apple Machine Learning Research
Apple Machine Learning Research
P
Proofpoint News Feed
量子位
U
Unit 42
T
The Blog of Author Tim Ferriss
Microsoft Azure Blog
Microsoft Azure Blog
V
V2EX
O
OpenAI News
S
Secure Thoughts
罗磊的独立博客
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Google Online Security Blog
Google Online Security Blog
Cloudbric
Cloudbric
W
WeLiveSecurity
IT之家
IT之家

Swift for Visual Studio Code comes to Open VSX Registry | InfoWorld

Notion courts developers with a platform for AI agents and workflow automation Using continuous purple teaming to protect fast-paced enterprise environments A better way to work with SQL Server AWS debuts Graviton-powered Redshift RG instances to cut analytics costs SAP’s AI promises last year? Most are still rolling out First look: Lemonade serves up local AI with limitations GitLab CEO sees developer tool bill increasing 100-fold Red Hat adds support for agentic AI development What’s new and exciting in JDK 26 Kill the loading spinner with local-first data and reactive SQL A networking revolution at AWS Tokenmaxxing is super dumb How to add AI to an existing product (without annoying users) Your AI doesn’t need another database What happens when engineering teams reorganize around AI agents Python isn’t always easy When cloud giants meddle in markets 12 model-level deep cuts to slash AI training costs The best new features in Python 3.15 Teradata launches platform for enterprise AI agents moving beyond pilots Three skills that matter when AI handles the coding MongoDB targets AI’s retrieval problem Building AI apps and agents with Microsoft Foundry Designing front-end systems for cloud failure No, AI won’t destroy software development jobs Diskless databases: What happens when storage isn’t the bottleneck Vibe coding or spec-driven development? The agentic AI distraction Vibe coding or spec-driven development? How to choose Cloud providers are blinded by agentic AI SAP to acquire data lakehouse vendor Dremio Small language models: Rethinking enterprise AI architecture Making AI work through eval hygiene Improving AI agents through better evaluations AI in the cloud is easy but expensive Running AI in the cloud is easy – and expensive Making AI work for databases Harness teams of agentic coders with Squad Harness teams of coding agents with Squad Oracle NetSuite announces AI coding skills for SuiteCloud developers Why it’s so hard to create stand-alone Python apps A new challenge for software product managers The hidden cost of front-end complexity GitHub shifts Copilot to usage-based billing, signaling a new cost model for enterprise AI tools OpenAI’s Symphony spec pushes coding agents from prompts to orchestration The front-end architecture trilemma: Reactivity vs. hypermedia vs. local-first apps Enterprise AI is missing the business core The best JavaScript certifications for getting hired Google begins putting the guardrails on agentic AI Why world models are AI’s next frontier Where to begin a cloud career Google pitches Agentic Data Cloud to help enterprises turn data into context for AI agents How open source ideals must expand for AI Is your Node.js project really secure? How I doubled my GPU efficiency without buying a single new card SpaceX secures option to acquire AI coding startup Cursor for $60B Google’s Gemma 4 shines on local systems – both big and small AI is upending the SaaS game How AI is upending SaaS tools Snowflake offers help to users and builders of AI agents From the engine room to the bridge: What the modern leadership shift means for architects like me Addressing the challenges of unstructured data governance for AI The cookbook for safe, powerful agents Enterprises are rethinking Kubernetes GitHub pauses new Copilot sign-ups as agentic AI strains infrastructure Best practices for building agentic systems Making agents dull Oracle delivers semantic search without LLMs When cloud giants neglect resilience Exciting Python features are on the way Ease into Azure Kubernetes Application Network The agent tier: Rethinking runtime architecture for context-driven enterprise workflows The two-pass compiler is back – this time, it’s fixing AI code generation MuleSoft Agent Fabric adds new ways to keep AI agents in line Salesforce launches Headless 360 to support agent‑first enterprise workflows Tap into the AI APIs of Google Chrome and Microsoft Edge Where will developer wisdom come from? GitHub adds Stacked PRs to speed complex code reviews The hyperscalers are pricing themselves out of AI workloads HTMX 4.0: Hypermedia finds a new gear Google Cloud introduces QueryData to help AI agents create reliable database queries Hands-on with the Google Agent Development Kit Are AI certifications worth the investment? AWS targets AI agent sprawl with new Bedrock Agent Registry Cloud degrees are moving online Swift for Visual Studio Code comes to Open VSX Registry AI agents aren't failing. The coordination layer is failing How Agile practices ensure quality in GenAI-assisted development Anthropic rolls out Claude Managed Agents Microsoft’s reauthentication snafu cuts off developers globally Meta’s Muse Spark: a smaller, faster AI model for broad app deployment Bringing databases and Kubernetes together Rethinking Angular forms: A state-first perspective Minimus Welcomes Yael Nardi as CBO to Facilitate Strategic Growth Microsoft announces end of support for ASP.NET Core 2.3 Get started with Python’s new frozendict type AWS turns its S3 storage service into a file system for AI agents Microsoft’s new Agent Governance Toolkit targets top OWASP risks for AI agents The winners and losers of AI coding GitHub Copilot CLI adds Rubber Duck review agent
Why private AI is the smarter bet
David Linthicum · 2026-06-26 · via Swift for Visual Studio Code comes to Open VSX Registry | InfoWorld

Pricing models in the AI market won't stay the same forever. Rising token costs, security risks, and operational realities are driving AI back on-premises.

For the past several years, the default assumption in enterprise IT was that AI would follow the same path as many other workloads and settle into the public cloud. That assumption seemed reasonable on the surface. The hyperscalers had the infrastructure, GPU capacity, managed services, and developer ecosystems. If you wanted to move fast, public cloud AI looked like the obvious answer.

That logic is now being challenged by reality. As enterprises move from AI experiments to AI in production, they increasingly find that the public cloud is a convenient place to start but not the most practical place to stay. Enterprises are wondering if they can afford to base their long-term AI strategies on cost models they do not control, risks they cannot fully contain, and architectures that are optimized for provider scale rather than enterprise economics.

This is why private cloud AI is becoming more popular. Enterprises are not moving on-premises because it’s a fashionable choice. They are moving because, in many cases, it is the financially rational choice.

The expense of token-based AI

The market still treats token-based AI pricing as a stable, mature economic model. It is not. Much of what enterprises pay today reflects a highly competitive environment in which providers are still subsidizing adoption, offering aggressive discounts, and prioritizing market share over normalized margins. That may be good news in the short term, but it is dangerous to assume those conditions will persist.

As enterprises scale their usage, token consumption shifts from an interesting line item to serious financial exposure. A chatbot pilot is one thing. Enterprisewide inference across business operations, customer engagement, knowledge systems, automation, analytics, and embedded software is something else entirely. When AI becomes part of the daily operating fabric of the business, token charges stop being experimental expenses and become recurring utility bills. At that point, even modest changes in pricing can have major budget consequences.

Many tech leaders are now rethinking their assumptions about AI costs, realizing that current pricing may not reflect long-term expenses. As subsidies fade and usage increases, token costs are likely to rise sharply, potentially making large-scale public AI deployments less economically viable. That is the trap enterprises want to avoid. No CIO wants to explain that the company successfully operationalized AI only to discover that a growing bill from a public provider offsets every business gain. Enterprises have seen this before with cloud cost overruns, and they do not want to repeat it with AI.

Hybrid AI is the natural end state

It is becoming clear that the future of enterprise AI is neither all public cloud nor all on-premises. It is a hybrid. The market is maturing beyond ideology and moving toward workload placement based on economics, governance, latency, and control.

That shift matters because not every AI problem requires a giant hosted model. In fact, many enterprise use cases do not. A growing number of organizations are finding that smaller, domain-specific models can perform as well as, and often better than, larger ones for targeted business tasks. Some use tuned models. Some rely on classic machine learning and predictive systems. Some combine retrieval techniques with smaller language models. Others build tightly constrained models tailored to specific operational domains.

These systems are often better suited to private infrastructure. They run closer to enterprise data, can be optimized for predictable workloads, and avoid the open-ended cost profile of external tokenized services. This is especially true when the model is used repeatedly within internal business processes rather than occasionally by a limited set of users. In other words, enterprises are not just choosing private AI because they dislike public cloud pricing. They are choosing it because they are learning to build AI systems that meet enterprise requirements rather than defaulting to whatever is easiest to consume from the outside.

Security and governance

Cost may be the loudest concern, but it is not the only one. Security and governance are becoming equally powerful drivers. Enterprises are increasingly uncomfortable with the idea of sensitive information flowing through public AI tools, public APIs, and user workflows that are difficult to monitor and control. The concern is not abstract. Employees routinely paste confidential information into public AI interfaces to boost productivity. Development teams sometimes move faster than policy can keep pace. Business units adopt tools before governance can catch up. The result is a growing risk of data leakage, unauthorized exposure, compliance failures, and security incidents directly tied to the use of AI.

This changes the conversation. Once AI touches customer records, financial models, regulated data, or other proprietary information, the focus shifts from deployment speed to the risk you introduce to the core of the business. While public clouds can provide strong security, many enterprises prefer tighter internal controls for sensitive AI workloads to ensure better observability, access, data locality, and policy enforcement.

There’s no question that private AI reduces the number of unknowns. It gives enterprises more direct control over where data resides, how models are used, who can access them, and how systems are audited. That does not eliminate risk, but it makes risk easier to manage.

Private AI is harder but worth it

Private AI is not effortless. Building AI on premises or in a private cloud requires investment, planning, specialized skills, operational discipline, and a willingness to own more of the stack. Enterprises must think about infrastructure design, GPU utilization, life-cycle management, model operations, integration, and resilience in ways that public services often abstract away.

That extra work introduces real risk. Some organizations will underestimate the operational burden, some will overspend on infrastructure, and some will struggle to attract the right talent. Even with those challenges, many enterprises are concluding that the cost savings are too compelling to ignore.

Enterprises are not moving toward private AI because it is easier. They are moving because it’s smarter in the long term. They would rather take on more responsibility now than remain exposed to a pricing model that could become unsustainable later. They would rather invest in owned capability than rent critical intelligence from an outside platform with uncertain future economics.

The public cloud will remain important, especially for experimentation, bursting, and select services. But for many production workloads, the balance is shifting. As token costs rise, governance pressures intensify, and organizations become better at building focused models rather than defaulting to giant LLMs, more enterprises will conclude that their most valuable AI belongs closer to home.