Private AI Platforms: VMware VCF vs. Red Hat OpenShift AI

Private AI Platforms: VMware VCF vs. Red Hat OpenShift AI

One NVIDIA H100 can support roughly 50–80 concurrent engineers running LLM inference for code assistance — that figure comes from VMware’s own internal benchmarking, and it is the kind of number that turns a vague “we should do private AI” mandate into a rack-level capacity plan. Both incumbent infrastructure vendors now want to sell you that rack. Broadcom has folded Private AI Services into VMware Cloud Foundation 9.0 as a default component, and Red Hat has pushed OpenShift AI into serious inference territory with vLLM, the llm-d distributed inference project, and audited financial-services benchmarks.

The timing is not accidental. Surveys consistently rank data privacy as the top enterprise AI deployment concern — 72% of enterprises in one widely cited figure — and the EU AI Act’s high-risk obligations reach full enforcement in August 2026. If your board wants generative AI and your counsel wants the data on-premises, this is the comparison that decides your next infrastructure cycle. Here is the brief.

The verdict

If you are staying on VCF, use what you are already paying for. Private AI Services ships inside VCF 9.0 at no additional cost, and for a virtualization-first shop it is the shortest path from “no AI” to “governed AI” — weeks, not quarters. If your platform direction is Kubernetes-first, or the Broadcom renewal quote has you pricing exits, Red Hat OpenShift AI is the stronger strategic bet: it is built on vLLM, the open-source serving engine much of the industry has standardized on, and Red Hat has published audited benchmarks to back its inference claims. NVIDIA is the constant — both stacks are NVIDIA-validated, so your GPU investment transfers either way. The real question is where your operational center of gravity sits: vCenter or Kubernetes.

What changed

Broadcom made VCF 9.0 an AI-native platform. Private AI Services — the evolution of VMware Private AI Foundation with NVIDIA — became a standard part of the VCF subscription starting in the first quarter of 2026, at no extra cost. That bundle now covers a model store, model runtime, agent builder, vector database, data indexing and retrieval, and GPU monitoring. Hardware support has moved to the Blackwell generation: NVIDIA B200 and RTX PRO 6000 Blackwell Server Edition GPUs, plus ConnectX-7 and BlueField-3 DPUs.

Red Hat, meanwhile, rebuilt OpenShift AI’s serving layer around vLLM and launched llm-d, an open-source framework for distributed inference on Kubernetes that adds cache-aware intelligent routing and materially better P95/P99 latency at scale. More telling for buyers: Red Hat produced the first audited STAC-AI LANG6 inference results on a containerized Kubernetes platform, run with NVIDIA and Supermicro hardware against financial-services workloads. Audited third-party numbers are rare in this market. Takeaway: in about twelve months, private AI went from optional add-on SKU to default platform capability at both vendors.

Side by side

VCF Private AI ServicesRed Hat OpenShift AI
SubstratevSphere/VCF 9.0, VM-centric with Kubernetes on topOpenShift (Kubernetes) on bare metal, VMs, or cloud
Serving engineModel Runtime, NVIDIA NIM-alignedvLLM, plus llm-d for distributed inference
PackagingIncluded in VCF 9.0 subscription by defaultSeparate subscription on top of OpenShift
Published proofInternal sizing benchmarks (one H100 ≈ 50–80 concurrent engineers)Audited STAC-AI LANG6 results; MLPerf inference submissions
AcceleratorsNVIDIA-validated, Blackwell-generation supportNVIDIA-first; AMD and Intel accelerator options as of mid-2026
Best fitLarge vSphere estates staying with BroadcomKubernetes-first platform teams; VMware leavers

Where VCF Private AI Services is strong

Broadcom’s pitch is operational continuity, and it is genuinely persuasive for a vSphere-first shop. GPU-backed model workloads are managed like any other VCF workload — same lifecycle tooling, same admin skills, same DR posture your team already runs. The model store gives governance teams a curated, versioned catalog, which matters more than demos suggest once audit and the EU AI Act’s documentation requirements enter the picture. And the packaging change is the headline: because Private AI Services is included in VCF 9.0, the marginal software cost of a first AI workload on an existing estate is near zero. Your real spend is GPUs, power, and cooling — we have broken down that math in our on-prem GPU cluster cost guide.

Who it fits: enterprises with 70%+ of workloads on vSphere, VMware-trained operations teams, and a signed multi-year VCF agreement. If that is you, the fastest governed pilot you can run this quarter is the one you have already licensed.

Where OpenShift AI is strong

Red Hat’s advantage is upstream gravity. vLLM has become the de facto open-source inference engine, and Red Hat employs a significant share of its maintainers following the Neural Magic acquisition — meaning optimizations land in OpenShift AI close to where they are invented. llm-d extends that with disaggregated, cache-aware serving across GPU pools, which is what large-scale multi-tenant inference actually requires. The audited STAC-AI results give risk-averse buyers — banks especially — something no internal vendor benchmark can: third-party verification on realistic workloads.

Portability is the second argument. The same OpenShift AI platform runs on bare metal (no hypervisor layer on GPU nodes), on vSphere, or in public cloud. And for organizations exiting VMware over renewal economics, OpenShift Virtualization consolidates VMs, containers, and AI on one platform — we compared that path against the alternatives in our OpenShift Virtualization exit analysis. Who it fits: platform-engineering organizations already fluent in Kubernetes, and anyone whose three-year plan does not include Broadcom.

The honest downsides

VCF first. Private AI Services is young — several of its services shipped to general availability only this year, and the ecosystem around it is narrower than the Kubernetes AI ecosystem by an order of magnitude. It is also deeply coupled to NVIDIA; if you want accelerator optionality, this is not the stack. And the elephant: Broadcom’s per-core subscription model and licensing minimums have burned enough goodwill that “included at no extra cost” lands skeptically — the AI services are free, but the platform they require is not getting cheaper.

OpenShift AI has its own tax. It demands real Kubernetes and MLOps skills — teams without them ship their first model months later than they planned. Total cost stacks OpenShift plus the OpenShift AI subscription, and list pricing varies enough by deal size that you should model both platforms per GPU node, not per socket. Red Hat’s AI portfolio — RHEL AI, OpenShift AI, AI Inference Server — also takes genuine effort to map onto a buying decision. And a caution that applies to both: every dollar here concentrates risk on NVIDIA supply, pricing, and the separately licensed NVIDIA AI Enterprise layer. Neither vendor shields you from that.

What to do about it

  • VCF shops staying put: pilot Private AI Services on VCF 9.0 this quarter. It is already in your entitlement; the only new line items are GPUs and NVIDIA AI Enterprise.
  • Kubernetes-first or exiting VMware: choose OpenShift AI and fold the pilot into your migration business case — one platform decision instead of two.
  • Sizing rule of thumb: plan on one H100-class GPU per 50–80 concurrent code-assist users; halve that for RAG-heavy workloads, and validate with your own prompts before the purchase order goes out.
  • Check the buy-vs-API crossover first: below sustained utilization, hosted APIs still win on cost — run the numbers in our self-hosted LLM vs. API cost analysis before committing capital.
  • EU AI Act deadline: August 2026 enforcement makes model provenance, logging, and documentation table stakes. Score both platforms’ governance tooling against your compliance checklist, not the vendor demo.

The repeatable line for your next steering meeting: the private AI platform decision is a platform-loyalty decision. Pick the stack your operations team can run at 2 a.m., because the models will change faster than the infrastructure underneath them.

Frequently asked questions

What is VMware Private AI Foundation with NVIDIA?

It is the joint Broadcom–NVIDIA architecture for running generative AI on VMware Cloud Foundation with NVIDIA GPUs and NVIDIA AI Enterprise software. In VCF 9.0 it has evolved into Private AI Services — model store, model runtime, agent builder, vector database, and GPU monitoring — included in the VCF subscription by default.

Is OpenShift AI better than VMware for private AI?

Neither is universally better. OpenShift AI leads on open-source inference technology (vLLM, llm-d) and audited benchmarks; VCF Private AI Services leads on operational simplicity for existing vSphere estates and on packaging, since it is bundled with VCF 9.0. Choose based on whether your team’s core skill set is Kubernetes or vSphere.

How many GPUs do I need to run a private LLM?

As a planning heuristic, VMware’s internal benchmarking showed a single H100 supporting 50–80 concurrent engineers on code-assist inference. A quantized 7–13B model fits on one modern data-center GPU; 70B-class models generally need multiple GPUs or a distributed serving layer such as llm-d. Benchmark with your own workload before buying.

Do both platforms require NVIDIA GPUs?

VCF Private AI Services is an NVIDIA-specific stack. OpenShift AI is NVIDIA-first but also supports AMD and Intel accelerators as of mid-2026, which gives it more hardware optionality if GPU supply or pricing becomes a constraint.

Does the EU AI Act require on-premises AI?

No. The Act mandates governance, documentation, and risk controls — not where models run. On-premises deployment can simplify data-residency and confidentiality arguments, which is why enforcement beginning in August 2026 is accelerating private AI projects, but a well-governed cloud deployment can also comply.

Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.