惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Jina AI
Jina AI
云风的 BLOG
云风的 BLOG
人人都是产品经理
人人都是产品经理
T
The Blog of Author Tim Ferriss
阮一峰的网络日志
阮一峰的网络日志
罗磊的独立博客
J
Java Code Geeks
博客园 - 聂微东
B
Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
WordPress大学
WordPress大学
腾讯CDC
L
LangChain Blog
Apple Machine Learning Research
Apple Machine Learning Research
Microsoft Azure Blog
Microsoft Azure Blog
D
DataBreaches.Net
The GitHub Blog
The GitHub Blog
美团技术团队
博客园 - Franky
Google DeepMind News
Google DeepMind News
V
V2EX
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
月光博客
月光博客
The Cloudflare Blog

Help Net Security

Police arrest 10 suspected members of Black Axe cybercrime gang ShinyHunters claims it stole 1.4 million records from Udemy Sevii unveils Cyber Swarm Defense Mode to stop AI-driven attacks at scale Alleged Chinese hacker extradited to US over cyberattacks targeting COVID-19 research Cequence Agent Personas bring granular control and governance to enterprise AI agents NowSecure MARI gives enterprises evidence-based visibility into third-party mobile app risk The metrics killing your SOC, and what to use instead US state privacy fines reached $3.425 billion in 2025 Canada’s first SMS blaster case leads to three arrests Linux storage management tool Stratis 3.9.0 adds online encryption and cache-less pool startup TLS Connect gives SMBs a right-sized automated tool to manage TLS certificates Aptori expands its platform with autonomous offensive testing to reduce security bottlenecks Your IAM was built for humans, AI agents don’t care The AI criminal mastermind is already hiring on gig platforms 25 open-source cybersecurity tools that don’t care about your budget Product showcase: LuLu reveals unauthorized outbound connections from Mac apps Week in review: Claude Mythos finds 271 Firefox flaws, Vercel breach Users advised to drop passwords and make room for passkeys - Help Net Security Indirect prompt injection is taking hold in the wild - Help Net Security Compromised everyday devices power Chinese cyber espionage operations - Help Net Security New Cisco firewall malware can only be killed by pulling the plug - Help Net Security Meta is overhauling how you sign in, manage settings, and protect your accounts - Help Net Security Ubuntu 26.04 LTS delivers memory-safe system tools and live patching for Arm servers - Help Net Security OpenAI’s GPT-5.5 is out with expanded cybersecurity safeguards - Help Net Security AI is speeding up nation-state cyber programs - Help Net Security A study of 1,000 Android apps finds a privacy policy logging gap - Help Net Security IT spending to hit $6.31 trillion record, thanks to AI - Help Net Security Where AI in CI/CD is working for engineering teams - Help Net Security With AI's help, North Korean hackers stumbled into a near-undetectable attack - Help Net Security Hacker with a special interest in breaching sports institutions ends behind bars - Help Net Security
Multi-model AI is creating a routing headache for enterpr...
Anamarija Po · 2026-05-07 · via Help Net Security

Application teams are moving AI inference into production systems that support business operations. Enterprises are expanding traffic management, identity controls, observability, and routing systems for multiple AI models and environments.

F5’s 2026 State of Application Strategy Report found that 78% of organizations operate their own inference services and 77% identify inference as their primary AI activity. They also operate or evaluate an average of seven AI models.

AI inference is the process of using a trained AI model to generate responses, predictions, or decisions from new data.

AI inference operations

Inference moves into enterprise operations

AI inference falls into the same operational category as other enterprise application workloads. Teams run inference across public cloud platforms, colocation facilities, and on-premises infrastructure using many of the same controls tied to application delivery and security operations.

Multi-model AI inferencing introduces the same architectural and security challenges associated with distributed production workloads. Inference deployments are growing, and operational contro

“AI inference is becoming core to the business, which means AI delivery is now a traffic management challenge, and AI security is now a governance and control challenge. The companies that understand this shift early will be the ones that move faster and more safely,” said Kunal Anand, Chief Product Officer at F5.

New inference responsibilities are creating new teams with their own preferred tools, and the way firms manage the resulting complexity could shape the outcomes of their AI deployments.

Companies that underestimate infrastructure demands, complexity, and security risks tied to AI inference may encounter higher costs and operational strain.

Cross-model observability, centralized controls, and shared protection systems are becoming part of multi-model AI operations across enterprise environments.

AI workloads expand across hybrid multicloud operations

Most firms run hybrid multicloud environments across their own data centers, colocation facilities, and public cloud providers. They are integrating inference into business systems within hybrid multicloud environments.

Companies are also modifying external-facing applications to interact with AI agents. They are implementing identity-aware infrastructure to route and manage traffic based on machine or agent identity. Some are developing public-facing APIs that allow AI agents to access application data and functions, while others are adopting semantic data standards and data enrichment practices to support contextual understanding inside AI systems.

AI systems are becoming part of operational automation. AI now participates in decision-support functions and operational execution tasks tied to application environments. People continue to oversee application security, compliance, and business-risk decisions throughout enterprise systems.

Enterprises manage inference for multiple AI models

Organizations rely on multiple models, which challenges the idea that inference is a single endpoint. Instead, they operate a portfolio of models and services. This reflects the fact that no single model satisfies every workload, and teams continue to evaluate different models for different use cases. Different models introduce different costs, interfaces, and failure patterns under load.

Firms select AI models based on business and technical requirements that include cost optimization, compliance, resiliency, API compatibility, and model-specific capabilities.

Enterprises manage inference traffic across multiple models to support availability, preserve existing integrations, and control operational costs.

The shift toward multi-model AI is driven primarily by operational and business requirements. Models serve operational roles tied to workload requirements. Some are optimized for general tasks, while others are designed for specific workloads, cost efficiency, throughput, or accuracy.

Organizations are managing inference within distributed systems environments. They must determine which model should handle each request based on API compatibility, latency, availability, security, compliance, and cost.

Control planes become central to AI operations

AI systematization has direct implications for organizational architecture. When enterprises distill large models into smaller ones, combine them, or chain them dynamically, attention changes toward the control plane that determines where inference traffic goes, why it goes there, and how it is protected.

Orchestrating multiple models turns inference into a managed workload subject to delivery, security, cost control, and resilience requirements. Designing and managing systems that govern how inference traffic is routed, constrained, secured, and observed is becoming a major architectural priority for enterprises treating inference as a new application tier.

AI delivery and security converge around inference

Organizations are using AI to improve decision-making and automate operational tasks within defined limits. They are managing systems, policies, and controls connected to inference workloads

Many of them coordinate multiple AI models and inference services to support availability, compliance, and operational requirements. Investment is increasing around delivery and security controls tied to inference traffic and prompt handling.