NVIDIA now publishes Enterprise Reference Architectures for three distinct classes of on-prem “AI factory” — and if you are budgeting GPU infrastructure for 2026–27, those three documents will shape your shortlist more than any RFP you write. The tiers run from RTX PRO servers for departmental inference, through HGX-based systems in the mid-range, up to rack-scale GB300 NVL72 machines with 72 Blackwell Ultra GPUs and 36 Grace CPUs per rack, stitched together with Spectrum-X networking and BlueField DPUs. Validated designs starting at roughly four-node clusters mean the entry point is no longer a hyperscaler-only conversation.
“AI factory” is a marketing term. Underneath it is a real architectural decision with a ten-year procurement shadow. This brief decodes the tiers into a sizing decision tree, explains why the full-stack packaging should remind you of converged infrastructure circa 2014, and covers where Broadcom/VMware and Red Hat fit — because the software layer you pick determines how locked in the hardware layer leaves you.
What changed
NVIDIA’s Enterprise Reference Architectures (ERAs) have matured from loose sizing guides into full validated designs, and as of mid-2026 they define three named AI factory configurations: the RTX PRO AI Factory built on RTX PRO servers, the HGX AI Factory built on NVIDIA-Certified HGX systems, and the NVL72 AI Factory built on rack-scale GB300 NVL72 platforms. Each document prescribes the compute nodes, the Spectrum-X Ethernet fabric, BlueField DPU placement, storage partner options, and the NVIDIA AI Enterprise software stack on top. OEMs — Dell, HPE, Lenovo, Supermicro, Cisco — ship systems certified against these designs rather than inventing their own.
The practical shift: an enterprise no longer buys GPUs; it buys a validated cluster with the network and software pre-decided. NVIDIA claims the GB300 NVL72 delivers up to 50× the AI factory output of Hopper-generation systems and 30× faster real-time inference on trillion-parameter models. Treat vendor multipliers with the usual skepticism — but the direction is not in dispute. Rack-scale is the new unit of purchase at the top end.
The three tiers, decoded
| RTX PRO AI Factory | HGX AI Factory | NVL72 AI Factory | |
|---|---|---|---|
| Compute unit | 2U RTX PRO servers, air-cooled | NVIDIA-Certified HGX nodes (8 GPUs/node class) | GB300 NVL72 rack: 72 Blackwell Ultra GPUs, 36 Grace CPUs |
| Entry scale | ~4-node clusters upward | Small multi-node clusters to hundreds of GPUs | One rack minimum; multi-rack pods |
| Primary workloads | Departmental inference, RAG, VDI-adjacent AI, fine-tuning small models | Serious fine-tuning, mid-size training, high-throughput inference | Foundation-model training, trillion-parameter reasoning, agentic pipelines |
| Facilities impact | Standard racks and power | High-density racks; liquid cooling increasingly assumed | Liquid cooling and 100kW+ rack power, full stop |
| Who it fits | Most enterprises starting private AI | Enterprises with committed AI product roadmaps | Model builders, sovereign AI, GPU service providers |
The honest read on each tier. The RTX PRO tier is the one most IT shops should study first — it runs on facilities you already have, and for inference-heavy private AI (which is what most enterprise AI actually is) it is usually enough. The HGX tier is the default answer when data science teams demand training capacity; it is also where cost overruns live, because the network and storage bill surprises people. The NVL72 tier is genuinely impressive engineering, but if you have to ask whether you need it, you do not. Run the arithmetic in our on-prem GPU cluster cost guide before any of these conversations.
Why it matters
We have seen this movie. A decade ago, converged and hyperconverged infrastructure took the server-network-storage decision away from component buyers and sold it as one validated SKU. It worked — deployment risk fell, time-to-production fell — and vendor optionality fell with it. NVIDIA’s ERAs do the same for AI: compute, Spectrum-X networking, BlueField DPUs, and NVIDIA AI Enterprise software arrive as a package, and every layer you accept is a layer you will not competitively bid later.
The networking clause is the one to read twice. The reference designs standardize on Spectrum-X Ethernet, which is excellent — and which quietly displaces the Arista, Cisco, or Juniper fabric your network team would otherwise have specified. Same with BlueField DPUs for east-west security and storage offload. None of this is bad engineering; all of it is deliberate account expansion. The takeaway for the meeting: validated designs buy you speed today at the price of negotiating leverage in 2029.
Where Broadcom/VMware fits
Broadcom’s answer is VMware Private AI Foundation with NVIDIA, which layers vGPU-partitioned NVIDIA AI Enterprise onto VMware Cloud Foundation. Its strength is real: if you are already a VCF shop, your operations team keeps the tooling it knows — vCenter, DRS, snapshots, the whole virtualization discipline — while data scientists get self-service GPU workstations and model runtimes. GPU sharing via vGPU is the underrated feature; departmental inference rarely saturates a full GPU, and virtualization claws that waste back.
The weaknesses are equally real. You are stacking two aggressive licensing regimes — Broadcom’s post-acquisition VCF subscription pricing plus NVIDIA AI Enterprise — on the same cluster, and Broadcom’s pricing changes since 2024 have made renewal math painful for mid-size shops. It fits large VMware-committed enterprises running mixed workloads; it fits poorly if you are trying to exit VCF or if your AI platform will be container-native from day one. Our VMware vs. OpenShift AI comparison works through that decision in detail.
Where Red Hat fits
Red Hat’s play is OpenShift AI: a Kubernetes-native MLOps platform where the NVIDIA GPU Operator manages device lifecycle cluster-wide and NVIDIA AI Enterprise is a certified overlay rather than the operating model. Its strength is portability — the same OpenShift AI stack runs on bare metal, on vSphere, in public cloud, or at the edge, which makes it the natural hedge against exactly the lock-in the ERAs encourage. Model serving, pipelines, and open source model support (including Red Hat’s vLLM-based inference work) are first-class, not bolted on.
The trade-off is operational: OpenShift assumes platform-engineering maturity. If your organization does not already run Kubernetes competently, standing up OpenShift AI alongside new GPU hardware is two transformations at once, and the failure mode is a very expensive science project. It fits container-fluent enterprises and regulated shops that need hybrid portability; it does not fit teams whose operational center of gravity is still the vSphere client.
What to do about it
A sizing decision tree you can defend in a budget meeting:
- Workload is inference, RAG, or fine-tunes under ~70B parameters: start at the RTX PRO tier, four to eight nodes. Do not let anyone sell you an HGX cluster for a chatbot.
- Sustained training or GPU utilization forecast above ~60% around the clock: the HGX tier beats cloud economics over a 3-year horizon; below that threshold, rent first and buy after you have utilization data.
- Training foundation models or serving trillion-parameter reasoning at scale: NVL72 territory — and a facilities project before it is an IT project. Budget the liquid cooling retrofit honestly.
- Negotiate the fabric separately. Accept the validated compute design, but make Spectrum-X win the network on price against at least one alternative bid, even if you expect it to win.
- Pick the software layer for your team, not the diagram. VMware-committed and virtualization-strong: Private AI Foundation. Kubernetes-strong and portability-minded: OpenShift AI. Neither: fix that before buying racks.
The one-line version for your VP: buy the smallest validated tier that covers the next 18 months, keep the network and software layers contestable, and let utilization data — not reference architecture diagrams — justify the next tier up.
Frequently asked questions
What is an NVIDIA AI factory?
It is NVIDIA’s term for a data center built as a production line for AI — GPU compute, high-speed networking, storage, and software packaged to turn data into models and tokens. In enterprise practice it means a cluster built to one of NVIDIA’s Enterprise Reference Architectures rather than a hand-assembled GPU farm.
What is the difference between GB300 NVL72 and HGX systems?
HGX systems are conventional servers with up to eight GPUs each, clustered over a network. GB300 NVL72 is a single liquid-cooled rack where 72 Blackwell Ultra GPUs and 36 Grace CPUs behave as one NVLink-connected accelerator — bought, powered, and operated as a rack-scale unit.
Do I need NVIDIA AI Enterprise to build an AI factory?
The reference architectures assume it, and both VMware Private AI Foundation and Red Hat OpenShift AI integrate it. You can run open source stacks on the same hardware, but you give up the certified support matrix — a trade most regulated enterprises decline.
Is on-prem AI infrastructure cheaper than cloud GPUs?
Only above a utilization threshold — as a rule of thumb, sustained utilization north of roughly 60% over three years favors owning, and spiky or exploratory workloads favor renting. Data gravity, sovereignty rules, and per-token inference volume shift the math case by case.
Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.
