<-- Back to All News

VMware AI Factory and other new AI innovations in VCF Announcements at Explore 2026

Publish Date: September 4, 2026
Executive Overview

The strategic agenda for enterprise artificial intelligence infrastructure is undergoing a fundamental paradigm shift: the rapid transition from isolated, conversational proofs-of-concept toward persistent, autonomous agentic artificial intelligence systems. Across global financial institutions, healthcare networks, defense contractors, and multinational manufacturing conglomerates, software engineering groups are deploying autonomous agents capable of iterative multi-turn reflection, dynamic sub-task decomposition, stateful multi-agent collaboration, and automated external tool execution. While agentic paradigms unlock unprecedented levels of business process automation, they simultaneously dismantle the architectural, economic, and security assumptions upon which enterprise private cloud infrastructure was historically constructed.

Unlike traditional single-turn inference requests characterized by predictable request-and-response transactional profiles, autonomous agentic workloads operate in non-linear, recurrent computational loops. These workloads aggressively bloat context windows, trigger unpredictable token consumption bursts, saturate high-speed fabric interconnects, and require continuous, low-latency access to distributed corporate knowledge bases. Enterprise technology executives attempting to operationalize agentic AI have historically been forced to navigate an acute operational trilemma:

  1. Operational Complexity: Modern AI infrastructure demands intricate coordination across physical bare-metal hardware discovery, firmware harmonization, PCIe switch topology tuning, Remote Direct Memory Access (RoCEv2/InfiniBand) lossless fabric routing, and container runtime orchestration. Traditionally, provisioning high-density GPU nodes took weeks to months, resulting in rigid, unmanaged infrastructure silos and substantial toolchain sprawl.
  2. Unpredictable Token Economics: Public cloud large language model (LLM) APIs charge on a per-token basis. As autonomous agents execute recursive internal reasoning loops, error-correction cycles, and multi-agent communications, token consumption escalates exponentially. Industry analysts project enterprise token expenditures to grow by up to 24-fold by 2030, exposing organizations utilizing public cloud APIs to volatile, unsustainable operational costs and punitive data egress fees.
  3. Severe Governance and Security Vacuums: Autonomous agents that generate dynamic code and execute programmatic actions against production APIs introduce significant operational hazards. Unchecked agents operating without deterministic boundary controls can inadvertently execute destructive actions—such as dropping database tables, deleting virtual machines, or leaking sensitive customer telemetry across network perimeters. Furthermore, transmitting proprietary corporate knowledge to multi-tenant public cloud endpoints introduces severe data sovereignty, regulatory compliance (GDPR, HIPAA, DORA), and intellectual property risks.

This technical infrastructure advisory provides a comprehensive architectural evaluation of the strategic disclosures presented by Shobhit Bhutani, Principal Product Marketing Manager and Generative AI Lead within the VMware Cloud Foundation (VCF) Division of Broadcom, regarding the formal launch of VMware AI Factory and the broader VMware Private AI Cloud suite unveiled at VMware Explore 2026. Engineered as the software-defined operational core of VMware Private AI Cloud, VMware AI Factory establishes a unified, secure, and cost-effective execution tier for enterprise AI. By pairing automated bare-metal-to-model deployment with multi-tenant model sharing, an intelligent AI Gateway, a sandboxed Secure Agent Framework, event-driven Model Autoscaling, official NVIDIA-Certified Hypervisor validation, and an expanded matrix of optimized open-weight models (including NVIDIA Nemotron 3, Google DeepMind Gemma 4, NEC cotomi, Alibaba Qwen3.8-27B, and Zhipu AI GLM 5.2), Broadcom delivers a definitive blueprint for running, governing, and scaling private enterprise AI on-premises.

Features

The technical enhancements introduced within VMware AI Factory and VMware Private AI Cloud establish an enterprise-grade capabilities matrix designed to eliminate manual Day-0 provisioning friction, maximize expensive hardware accelerator utilization, and enforce continuous, hardware-rooted zero-trust governance across the AI lifecycle.

  • VMware AI Factory Software-Defined Deployment Engine: Serving as the foundational software engine of VMware Private AI Cloud, VMware AI Factory compresses the end-to-end timeline from bare-metal server ingestion to active model serving from weeks down to hours. The engine coordinates automated physical hardware discovery, baseboard management controller (BMC) interrogation, ESXi hypervisor deployment, VMware vSAN Express Storage Architecture (ESA) allocation, software-defined network segmentation via VMware NSX Virtual Private Clouds (VPCs), and Kubernetes control plane instantiation through VMware vSphere Kubernetes Service (VKS). Through certified partnerships with server original equipment manufacturers (OEMs)—including Dell PowerEdge, Cisco, Lenovo, and Supermicro AI ReadyNodes—the framework establishes standardized, repeatable compute blocks.
  • Integrated Heterogeneous Bare-Metal Automation with MetalSoft: Addressing the operational fragmentation beneath the hypervisor, Broadcom has partnered with MetalSoft to embed native, heterogeneous bare-metal orchestration directly into the VCF management plane. The integration enables platform engineers to discover, baseline, configure BIOS/BMC firmware, and repave physical servers across multi-vendor fleets (Dell, HPE, Lenovo, Supermicro, and custom ODM hardware) directly from the VCF console. This reduces physical server preparation time from weeks to minutes, unifying software and hardware lifecycle management into a singular operational model.
  • Heterogeneous Hardware and Silicon Acceleration Architecture: VMware AI Factory decouples the upper artificial intelligence application tier from the underlying physical compute silicon, providing flexible execution across diverse accelerators:
    • NVIDIA Blackwell and Hopper Architectures: Leveraging the official NVIDIA-Certified Hypervisor status achieved by VMware vSphere 9.1 and future vSphere 9 releases, the platform delivers near bare-metal execution efficiency (less than 3% virtualization overhead) across Hopper (H100, H200) and Blackwell (B100, B200) platforms for GPU collective communication (NCCL), memory bandwidth, and distributed matrix calculations.
    • AMD Instinct Accelerators and Open ROCm Software Stack: In collaboration with AMD, the platform provides zero-touch provisioning across AMD Instinct GPU clusters. The architecture automates the deployment of the entire stack—from vSphere and vSAN through Kubernetes and the AMD GPU operator—utilizing the AMD DVX driver to attach physical GPUs directly to large virtual machines consumed by VKS clusters.
  • Multi-Tenant Model Sharing with Strict Namespace Isolation: To resolve the financial and operational waste of deploying dedicated, per-team model runtimes, VMware AI Factory enhances the Model Runtime service to enable secure model sharing across multiple enterprise tenants and business units. A single model runtime instance scales foundational models across shared GPU clusters, while independent application teams interact through isolated vSphere Namespaces. Dedicated tenant encryption keys, role-based access control (RBAC), and network micro-segmentation ensure that fine-tuning datasets, prompt histories, and retrieval-augmented generation (RAG) contexts remain strictly segregated across organizational boundaries.
  • Intelligent AI Gateway and Tokenomic Governance Engine: Designed to balance token cost optimization, operational latency, and domain specialization, the upcoming AI Gateway functions as an intelligent reverse proxy and traffic management layer for enterprise prompts:
    • Dynamic Prompt Routing: Intelligently inspects and steers incoming agent queries between locally hosted open-weight models on VCF and external public hyperscaler frontier models, optimizing for latency targets, reasoning complexity, and per-token cost boundaries.
    • Usage and Token Limiting: Enforces user-, department-, and application-level token quotas to prevent runaway loops or budget overruns during intensive multi-agent workflows.
    • Application Authorization: Validates incoming application identities prior to prompt routing, enforcing strict OpenID Connect (OIDC) token-based authentication and authorization policies.
  • Secure Agent Framework and Sandboxed Runtime Harness: Recognizing that autonomous agents generate and execute code dynamically, the platform incorporates a dual-layer containment architecture to prevent unauthorized infrastructure tampering:
    • Sandboxed Runtime Containment: Establishes a secure, virtualized container sandbox where dynamic, agent-generated code is executed in complete isolation from core virtual machines, storage datastores, and production database networks.
    • Agent Harness Governance Layer: Functions as the operational policy gatekeeper, regulating agent invocation sequences, defining explicit tool-calling access parameters, governing inter-agent communication protocols, and mathematically validating agent outputs before programmatic commands are executed against infrastructure.
  • Event-Driven Model Autoscaling: AI workloads cannot effectively accommodate erratic, bursty traffic profiles without automated capacity management. The platform introduces event-driven Model Autoscaling, allowing platform operators to establish deterministic thresholds for token latency and concurrent active sessions. When agentic activity causes latency or concurrency to breach pre-set boundaries, the platform automatically provisions additional model replicas and allocates GPU resources to maintain application service-level agreements (SLAs). Conversely, when demand subsides, instances are drained and decommissioned, returning GPU capacity to the shared cluster pool to prevent idle resource wastage.
  • Comprehensive Model Runtime Matrix Powered by vLLM: Utilizing vLLM as its default accelerated model runtime, VMware AI Factory delivers performance-optimized, out-of-the-box execution for more than 150 open-source and open-weight model architectures. Broadcom officially validated the following frontier model families on VCF:
    • NVIDIA Nemotron 3: Open, multimodal model family featuring a hybrid Mamba-Transformer Mixture-of-Experts (MoE) architecture, 1-million-token context windows, and multi-environment reinforcement learning optimized for long-running agent workflows.
    • Google DeepMind Gemma 4: Lightweight, open-weight multimodal model family engineered for high-throughput local execution and autonomous agent deployment.
    • NEC cotomi: Specialized model optimized for the Japanese language, trained on curated enterprise datasets to deliver high-speed processing with a 40% improvement in token efficiency.
    • Alibaba Qwen3.8-27B: Dense 27-billion-parameter vision-language model engineered for complex code refactoring, multimodal reasoning, and long-horizon tasks with togglable reasoning depth.
    • Zhipu AI GLM 5.2: Open-source general language model designed for local deployment of autonomous coding and multi-step reasoning agents under strict data sovereignty constraints.
  • Strategic Enterprise Ecosystem Alliances: Expanding the operational surface of Private AI Services, Broadcom integrates key technology partners into the platform:
    • Appian: End-to-end AI process automation for mission-critical enterprise and public sector workflows.
    • ClearML: Comprehensive AI orchestration layer governing GPU allocations, model registries, and agent lifecycles to enable internal AI-as-a-Service (AIaaS).
    • Eve Security: Runtime agent governance providing “agent-in-the-loop” behavioral monitoring, context-aware anomaly detection, and real-time Data Loss Prevention (DLP) enforcement.
    • Solo.io: Open-source agentic networking leveraging the kagent runtime and agentgateway data plane to govern inter-agent and tool-calling traffic on private infrastructure.
    • TrueFoundry: Enterprise-grade unified gateway encompassing LLM, Model Context Protocol (MCP), and Agent Gateway orchestration across distributed infrastructure.
Benefits

Operationalizing VMware AI Factory and VMware Private AI Cloud on VMware Cloud Foundation delivers measurable strategic, economic, and architectural advantages over conventional bare-metal clusters and unmanaged public cloud AI services.

  • Dramatic Compression of Time-to-Value for Accelerated Computing: Traditional data center workflows require weeks of manual coordination across server procurement, BIOS configuration, network cabling, storage mapping, and manual Kubernetes orchestration before data science teams can deploy an initial model. VMware AI Factory compresses this onboarding cycle to a few hours through automated, policy-driven orchestration spanning bare-metal bare metal provisioning, hypervisor installation, and workload domain instantiation, drastically accelerating enterprise innovation velocity.
  • Sustainable Tokenomics and Predictable Fixed Infrastructure Economics: Routing recursive agentic reasoning loops through public hyperscaler APIs exposes organizations to extreme financial volatility as multi-turn reflection and tool invocations multiply token volumes. By hosting high-concurrency model execution on private VCF infrastructure governed by the AI Gateway’s prompt routing and token-limiting policies, organizations establish fixed, predictable operational expenditures and avoid punitive hyperscaler consumption pricing.
  • Elimination of Hardware Fragmentation and GPU Capacity Waste: Dedicating isolated bare-metal servers to individual AI project teams inevitably results in depressed average hardware utilization, with unmanaged clusters frequently languishing below 30% utilization. Multi-tenant model sharing, dynamic fractional GPU slicing, and automated model autoscaling allow enterprises to consolidate multiple business units onto shared physical clusters, driving hardware utilization rates above 75% and significantly reducing the total cost of ownership (TCO) per inference token.
  • Hardware-Rooted Containment of Autonomous Agentic Execution Risks: Autonomous agents that generate arbitrary code or execute multi-system API calls pose profound operational hazards if permitted to interact directly with production environments. The Secure Agent Framework’s virtualized sandboxing and policy-governed Agent Harness establish deterministic control boundaries, ensuring that runaway agentic execution loops cannot compromise system integrity, corrupt underlying enterprise databases, or initiate unauthorized lateral network traversal.
  • Absolute Data Sovereignty, Privacy, and Regulatory Alignment: Hosting sensitive foundational models, fine-tuning datasets, vector databases, and real-time inference telemetry entirely within private cloud boundaries guarantees full compliance with strict regulatory mandates under GDPR, HIPAA, DORA, and regional sovereign cloud guidelines. Proprietary intellectual property, customer financial records, and healthcare telemetry remain permanently protected behind software-defined NSX micro-segmentation perimeters without exposure to external multi-tenant public cloud networks.
  • Enterprise Operational Continuity and Unified Fleet Governance: Unlike bare-metal GPU clusters that require bespoke operational toolchains, VCF Private AI environments integrate natively into standard enterprise IT operations. Virtualized GPU workloads leverage vSphere High Availability, Distributed Resource Scheduler (DRS), non-disruptive vMotion live migrations, and centralized SDDC Manager lifecycle management, allowing IT organizations to govern AI infrastructure using their existing staff competencies and operational playbooks.
Use Cases

Global enterprise organizations operating across heavily regulated, high-transaction, and data-intensive industry sectors can implement VMware AI Factory on VCF to address complex operational and architectural challenges.

  • Global Financial Services Sovereign Agentic Fraud Analytics and Trading: A multinational investment banking institution deploys an on-premises private AI factory powered by certified Dell and Supermicro server nodes equipped with NVIDIA Blackwell GPUs. The bank utilizes the Secure Agent Framework and VMware Data Services Manager to run autonomous fraud detection agents against real-time transactional ledgers and localized PostgreSQL vector databases. The AI Gateway routes latency-sensitive queries to local Qwen3.8-27B and GLM 5.2 model instances, keeping all transactional records within private data center boundaries. By leveraging multi-tenant model sharing, the institution provides independent compliance, algorithmic trading, and retail banking squads with isolated access to the shared inference tier, reducing infrastructure capital costs by 45% while adhering strictly to international data sovereignty frameworks.
  • Sovereign Public Sector Administration and Multi-Agency Digital Transformation: A national sovereign cloud agency provides shared infrastructure services to multiple government ministries and defense contractors. The agency implements VMware AI Factory integrated with MetalSoft bare-metal automation to rapidly onboard heterogeneous server fleets across regional sovereign data centers. Using localized instances of NEC’s cotomi and Google DeepMind’s Gemma 4, the agency hosts policy-governed document classification and citizen service agents. Multi-tenant model sharing guarantees that municipal agencies share underlying physical GPU capacity while retaining cryptographic namespace isolation, enabling rapid public sector digital transformation with zero data leakage across departmental jurisdictions.
  • Nationwide Healthcare Diagnostic Retrieval-Augmented Generation (RAG) and Clinical Agents: A nationwide healthcare provider consolidates distributed clinical research, pathology imaging, and patient portal assistants onto VMware Cloud Foundation. The clinical engineering team deploys NVIDIA Nemotron 3 models within VMware AI Factory to process complex multimodal diagnostic imaging datasets alongside patient Electronic Health Record (EHR) histories across 1-million-token context windows. The Agent Harness strictly controls agent tool invocation, ensuring diagnostic agents can query clinical databases but cannot alter primary medical records without physician authorization. Native vSAN ESA storage integration provides continuous, low-latency data streaming to GPU memory pools, ensuring real-time diagnostic recommendations without exposing protected health information (PHI) to third-party public cloud endpoints.
  • Large-Scale Retail Omnichannel Supply Chain Optimization and Agent Swarms: A global retail conglomerate manages regional inventory routing, dynamic pricing models, and vendor negotiations using distributed autonomous agents. During major seasonal shopping events, agent reasoning loops spike unpredictably as hundreds of distribution nodes recalculate inventory allocations concurrently. Using VMware AI Factory’s Model Autoscaling, the private cloud automatically scales model serving instances across AMD Instinct GPU clusters in response to active session queues, maintaining sub-second API responsiveness across distribution centers. When sales volume stabilizes, the platform automatically consolidates model instances, liberating compute capacity for batch predictive analytics and lowering overall per-token operational expenditures.
Alternatives

A thorough architectural evaluation requires comparing Broadcom’s VMware AI Factory and Private AI Cloud framework against alternative enterprise artificial intelligence delivery models.

  • Public Cloud Hyperscaler AI Monoculture (AWS Bedrock / Azure AI Foundry / Google Vertex AI): Under this model, organizations bypass on-premises infrastructure entirely, routing all enterprise AI workloads through managed public cloud AI services and proprietary frontier APIs. While public hyperscalers provide rapid initial prototyping and eliminate data center hardware management, this strategy introduces severe long-term financial liabilities driven by non-linear token pricing, continuous data egress fees, and API rate limiting. More critically, transmitting proprietary corporate knowledge bases and sensitive customer records to multi-tenant public cloud environments introduces severe data sovereignty, regulatory, and intellectual property exposure risks.
  • Bespoke DIY Bare-Metal Open-Source AI Infrastructure (Slurm / Kubernetes on Raw Metal): In this approach, internal platform engineering and data science teams assemble custom AI infrastructure using bare-metal servers, open-source Linux distributions, manual GPU driver installations, and raw Kubernetes or Slurm job schedulers. While bare metal maximizes low-level hardware control and avoids commercial virtualization licensing fees, it imposes an extraordinary operational burden on internal IT teams. Platform engineers must manually configure, secure, and maintain complex RDMA network fabrics, storage plugins, and container environments, resulting in delayed project lifecycles, low aggregate hardware utilization, and rigid operational silos that cannot interoperate with broader enterprise IT estates.
  • Isolated Single-Vendor Turnkey AI Appliances (Proprietary AI Hardware Stacks): Under this operational strategy, enterprises procure pre-integrated hardware appliances bundled with proprietary, vendor-locked software stacks and dedicated orchestration tools. While turnkey appliances simplify initial Day-0 deployment for isolated research groups, they create physical and logical operational islands within the enterprise data center. These appliances cannot easily share underlying compute, memory, networking, or storage pools with standard enterprise virtual machines or general-purpose container workloads, creating stranded hardware capacity, inflating capital expenditures, and fragmenting corporate identity, security, and governance policies.
  • Decentralized In-House VM Deployments (Uncoordinated Ad-Hoc Virtualization): In this legacy operational model, individual software engineering teams deploy custom virtual machines running ad-hoc AI model runtimes across general-purpose hypervisors without centralized orchestration or GPU acceleration policies. While this model leverages existing virtualization tooling, it lacks intelligent prompt routing, model autoscaling, unified token tracking, and hardware-accelerated runtime optimization. Decentralized VM deployments lead to massive GPU resource fragmentation, severe memory contention, unmonitored security exposures, and an inability to support stateful, multi-turn agentic reasoning loops at scale.
Alternative Perspective

While the capabilities introduced in VMware AI Factory and VMware Private AI Cloud deliver compelling architectural, operational, and financial advantages for enterprise AI, an objective technical analysis reveals critical operational prerequisites, architectural dependencies, and governance trade-offs that enterprise technology leadership must evaluate prior to adoption.

A primary technical consideration centers on the physical data center facility demands imposed by high-density accelerated computing hardware. While VMware AI Factory streamlines software and hypervisor orchestration, running modern GPU platforms—such as NVIDIA HGX Blackwell or AMD Instinct clusters—requires immense electrical power delivery and specialized thermal management. Dense 8-GPU server nodes frequently demand between 40 kW and 100 kW per rack, necessitating advanced direct-to-chip liquid cooling loops, dedicated coolant distribution units (CDUs), and substantial facility retrofits. Organizations operating legacy enterprise data centers designed exclusively for low-density air cooling must carefully evaluate the substantial capital expenditures and construction lead times required to upgrade facilities before they can deploy high-density AI Factory compute blocks.

Furthermore, realizing the strategic benefits of the upcoming AI Gateway, Model Autoscaling, and Secure Agent Framework introduces organizational and architectural governance hurdles. Implementing event-driven model autoscaling requires platform engineering teams to define precise latency and concurrency thresholds; improper calibration can trigger rapid resource churn, thrashing GPU allocations, and disrupting neighboring transactional workloads running on the same private cloud fabric. Similarly, establishing the Secure Agent Framework demands that security and application teams collaborate closely to define granular tool-calling permissions and API access boundaries. Treating agent sandboxing as a purely automated infrastructure feature without defining rigorous application-level authorization policies risks introducing operational friction or failing to contain complex logical errors generated by autonomous agents.

Finally, enterprise technology leaders must assess the long-term vendor alignment inherent in standardizing on Broadcom’s full-stack private cloud ecosystem. While unifying bare-metal provisioning, hypervisor virtualization, software-defined networking, container orchestration, and AI model runtimes simplifies operations and provides a single support hierarchy, it binds an enterprise’s AI roadmap directly to Broadcom’s core subscription licensing models and product engineering cadences. Platform architects must ensure that this architectural consolidation aligns with corporate multi-cloud strategies and that internal teams develop the competencies required to govern hybrid prompt routing across both private cloud clusters and external cloud endpoints.

Final Thoughts

The emergence of autonomous agentic artificial intelligence represents a permanent transformation in enterprise computing paradigms. The traditional practice of isolating AI experimentation within specialized bare-metal islands or relying entirely on public cloud APIs is no longer viable in an era characterized by exponential token growth, strict regulatory oversight, and complex multi-turn reasoning workflows. Broadcom’s VMware AI Factory and VMware Private AI Cloud deliver an enterprise-grade architectural blueprint that bridges the historical divide between accelerated AI computing and private cloud governance.

By integrating automated bare-metal-to-model deployment, multi-tenant model sharing, intelligent tokenomics governance via the AI Gateway, deterministic execution containment through the Secure Agent Framework, and broad model interoperability via vLLM, VMware Cloud Foundation establishes a unified private cloud execution engine. Enterprise technology leadership should capitalize on this architectural milestone by executing infrastructure power and cooling assessments, defining collaborative platform engineering practices uniting data science and operations teams, and establishing policy-driven AI Factory pilots to drive secure, scalable business innovation across the enterprise private cloud estate.

Source

Explore 2026: VMware AI Factory and other new AI innovations in VCF