David Davis

  • Shadow AI: The $670K Breach Premium and What to Do

    Shadow AI: The $670K Breach Premium and What to Do

    Shadow AI finally has a price tag. IBM’s 2025 Cost of a Data Breach Report found that breaches involving high levels of unsanctioned AI use cost an average of $670,000 more than breaches without it — roughly $4.63 million against a $3.96 million global average. One in five breaches studied involved shadow AI. If you have been trying to get budget for GenAI controls by waving vaguely at “data leakage,” you now have a number a CFO will sit still for.

    This brief covers what the IBM data actually says, why the ban-and-block reflex fails in practice, the three-layer control pattern that holds up — sanctioned alternatives, inline data protection, and training — and an acceptable-use policy skeleton you can adapt this week.

    What changed

    Until this report cycle, shadow AI risk was a governance talking point without a dollar figure. IBM’s data changed that. Beyond the $670K premium, the supporting numbers are the real indictment: 97% of organizations that suffered an AI-related breach lacked proper AI access controls, and 63% either had no AI governance policy or were still drafting one when the incident hit. Shadow AI showed up in 20% of breaches — more often than incidents involving sanctioned AI systems, at 13%.

    Read that again: the unsanctioned tools your employees are quietly using are now a more common breach factor than the AI you actually deployed. The takeaway for the staff meeting — shadow AI is no longer a policy hygiene item, it is a quantified breach multiplier.

    Why it matters

    Shadow AI breachesGlobal average
    Average breach cost~$4.63M~$3.96M
    Share of studied breaches20%13% (sanctioned AI)
    Customer PII exposed65%53%
    Data spanning multiple environments62%

    The pattern behind the premium is mundane. An employee pastes customer records, source code, or deal terms into a consumer chatbot on a personal account. There is no audit trail, no retention control, and no way to know what left the building — so when a breach occurs, scoping takes longer and disclosure obligations balloon. That is why shadow AI incidents skew so heavily toward PII exposure and multi-environment sprawl. For how these figures stack against overall breach economics, see our data breach cost benchmarks.

    Quick diagnostic: if you cannot name the five GenAI tools your workforce used most last week, assume you are in the 63% without a working policy — and budget accordingly.

    Why ban-and-block fails

    The reflex response is a firewall rule and a memo. It does not work, and the failure mode is predictable. GenAI delivers too much personal productivity for a prohibition to hold: employees route around blocks with personal phones, home machines, and personal accounts — exactly the channels you cannot see. A block without a sanctioned alternative does not reduce usage; it relocates usage to where your DLP is blind. You trade visible, governable risk for invisible risk, which is precisely the condition IBM’s data prices at a $670K premium.

    Rule of thumb: never block a GenAI category until the sanctioned replacement is live and demonstrably good enough that reaching for the shadow tool feels like extra work. Prohibition is a sequencing decision, not a strategy.

    The control stack that works

    Layer 1: sanctioned alternatives with real data boundaries

    The cheapest way to shrink shadow AI is to give people a tool they are allowed to use. For Microsoft shops, Microsoft 365 Copilot and Copilot Chat with enterprise data protection put prompts and responses under the same Data Protection Addendum terms as the rest of M365, with no training of foundation models on your data. The strength is ubiquity — it lives where the work already is. The honest caveat: Copilot inherits your M365 permission sprawl, so years of oversharing in SharePoint becomes instantly searchable, and per-seat licensing is a real line item at scale.

    Google’s equivalent is Gemini for Workspace, with enterprise-grade data protection, no model training on customer prompts, and compliance support that as of mid-2026 extends to frameworks like HIPAA and FedRAMP High. It is the natural pick for Workspace estates and is often simpler to reason about than Microsoft’s tiering. The gap: outside Workspace-first organizations, Gemini’s enterprise footprint and third-party integration depth still trail Microsoft’s. We compare the two head-to-head in our Copilot vs Gemini enterprise analysis.

    Layer 2: inline GenAI data protection at the SSE layer

    Sanctioned tools cover the approved path; the SSE layer covers everything else. Zscaler’s approach — AI app visibility and inline data protection running through its Zero Trust Exchange — is representative of the pattern: discover which AI apps are actually in use, then apply prompt-level DLP with block, isolate, or coach actions rather than a blunt deny. The coaching option matters; a just-in-time warning changes behavior in a way a connection reset never will. Zscaler’s strength is that it sees traffic regardless of which app is fashionable this quarter. Its limits are structural: traffic must traverse the proxy, so unmanaged devices and personal phones escape it, and DLP policies for prompts require genuine tuning effort before the alerts are trustworthy.

    Layer 3: the human control

    Most shadow AI exposure is a well-meaning employee, not a malicious one — which makes training a control, not a checkbox. KnowBe4 has extended its security awareness platform with AI-specific modules and its AIDA agents, which personalize training based on individual behavior. Relative to the rest of this stack it is inexpensive, and it targets the actual failure mode: someone who does not realize pasting a customer list into a chatbot is exfiltration. Be honest about its ceiling, though — KnowBe4-style training measurably reduces risky clicks but cannot enforce anything. It complements the technical layers; it never substitutes for them.

    What it doesRepresentative vendorWeak spot
    Sanctioned alternativeGoverned GenAI where work happensMicrosoft 365 Copilot, Gemini for WorkspaceLicense cost; permission sprawl
    Inline SSE protectionDiscover apps, prompt-level DLP, coach/blockZscalerBlind to unmanaged devices
    Awareness trainingChanges user behavior at the point of pasteKnowBe4No enforcement power

    A shadow AI acceptable-use policy skeleton

    Most searchers want the template, so here it is. Keep the finished document to roughly one page — policies nobody reads are how you end up in IBM’s 63%. Adapt these eight clauses:

    • Scope. Applies to all staff and contractors, on any device, using any AI tool with company data.
    • Approved tools. A named, versioned list (e.g., Copilot, Gemini for Workspace) — everything else requires the exception process.
    • Prohibited inputs. Customer PII, credentials, source code, unreleased financials, and anything under NDA never enter unapproved tools.
    • Account rules. Corporate SSO only; personal accounts on AI services are out of scope for company data, full stop.
    • Output handling. AI output gets human review before it ships to a customer, a regulator, or production.
    • New-tool path. A lightweight request process with a committed SLA — two weeks or less, or people will not use it.
    • Monitoring disclosure. State plainly that AI traffic on corporate networks and devices is logged and inspected.
    • Enforcement and exceptions. Graduated consequences, plus a real exception route so edge cases surface instead of hiding.

    Wire the policy into a broader governance structure — ownership, risk review, model inventory — using our enterprise AI governance frameworks guide.

    What to do about it

    Sequence it as 30/60/90. First 30 days: discover — pull AI app usage from your SSE or firewall logs, scan expense reports for AI subscriptions, and run an amnesty survey so people admit what they use without fear. Days 30–60: decide the sanctioned set, publish the one-page policy, and turn on enterprise data protections in whichever suite you already own. Days 60–90: enable inline GenAI DLP in coach mode before block mode, launch AI-specific awareness training, and report shadow-usage trend lines to the risk committee monthly.

    The sentence to repeat upstairs: shadow AI adds about $670K to a breach we would already struggle to afford, and the fix is mostly tools we already license plus one policy page. That argument wins budget meetings.

    Frequently asked questions

    What is shadow AI risk?

    Shadow AI risk is the exposure created when employees use AI tools that IT has not approved or configured — typically consumer chatbots on personal accounts. Because the usage is invisible, sensitive data can leave the organization with no audit trail, and IBM’s 2025 data shows such breaches cost about $670K more than average.

    Should companies block ChatGPT and other AI tools at work?

    Blanket blocking usually backfires — employees shift to personal devices where no control applies. The stronger pattern is a sanctioned alternative with enterprise data protection, inline DLP that coaches rather than just denies, and a clear acceptable-use policy.

    What should a shadow AI policy include?

    At minimum: scope, a named approved-tools list, prohibited data classes, corporate-account requirements, output review rules, a fast path for requesting new tools, monitoring disclosure, and enforcement with an exception process. One page is the right length.

    How do you detect shadow AI use in the enterprise?

    Start with SSE or secure web gateway logs (Zscaler and its peers categorize AI apps out of the box), review expense reports for AI subscriptions, and run an amnesty survey. Discovery before enforcement — you cannot govern what you have not counted.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Private AI Platforms: VMware VCF vs. Red Hat OpenShift AI

    Private AI Platforms: VMware VCF vs. Red Hat OpenShift AI

    One NVIDIA H100 can support roughly 50–80 concurrent engineers running LLM inference for code assistance — that figure comes from VMware’s own internal benchmarking, and it is the kind of number that turns a vague “we should do private AI” mandate into a rack-level capacity plan. Both incumbent infrastructure vendors now want to sell you that rack. Broadcom has folded Private AI Services into VMware Cloud Foundation 9.0 as a default component, and Red Hat has pushed OpenShift AI into serious inference territory with vLLM, the llm-d distributed inference project, and audited financial-services benchmarks.

    The timing is not accidental. Surveys consistently rank data privacy as the top enterprise AI deployment concern — 72% of enterprises in one widely cited figure — and the EU AI Act’s high-risk obligations reach full enforcement in August 2026. If your board wants generative AI and your counsel wants the data on-premises, this is the comparison that decides your next infrastructure cycle. Here is the brief.

    The verdict

    If you are staying on VCF, use what you are already paying for. Private AI Services ships inside VCF 9.0 at no additional cost, and for a virtualization-first shop it is the shortest path from “no AI” to “governed AI” — weeks, not quarters. If your platform direction is Kubernetes-first, or the Broadcom renewal quote has you pricing exits, Red Hat OpenShift AI is the stronger strategic bet: it is built on vLLM, the open-source serving engine much of the industry has standardized on, and Red Hat has published audited benchmarks to back its inference claims. NVIDIA is the constant — both stacks are NVIDIA-validated, so your GPU investment transfers either way. The real question is where your operational center of gravity sits: vCenter or Kubernetes.

    What changed

    Broadcom made VCF 9.0 an AI-native platform. Private AI Services — the evolution of VMware Private AI Foundation with NVIDIA — became a standard part of the VCF subscription starting in the first quarter of 2026, at no extra cost. That bundle now covers a model store, model runtime, agent builder, vector database, data indexing and retrieval, and GPU monitoring. Hardware support has moved to the Blackwell generation: NVIDIA B200 and RTX PRO 6000 Blackwell Server Edition GPUs, plus ConnectX-7 and BlueField-3 DPUs.

    Red Hat, meanwhile, rebuilt OpenShift AI’s serving layer around vLLM and launched llm-d, an open-source framework for distributed inference on Kubernetes that adds cache-aware intelligent routing and materially better P95/P99 latency at scale. More telling for buyers: Red Hat produced the first audited STAC-AI LANG6 inference results on a containerized Kubernetes platform, run with NVIDIA and Supermicro hardware against financial-services workloads. Audited third-party numbers are rare in this market. Takeaway: in about twelve months, private AI went from optional add-on SKU to default platform capability at both vendors.

    Side by side

    VCF Private AI ServicesRed Hat OpenShift AI
    SubstratevSphere/VCF 9.0, VM-centric with Kubernetes on topOpenShift (Kubernetes) on bare metal, VMs, or cloud
    Serving engineModel Runtime, NVIDIA NIM-alignedvLLM, plus llm-d for distributed inference
    PackagingIncluded in VCF 9.0 subscription by defaultSeparate subscription on top of OpenShift
    Published proofInternal sizing benchmarks (one H100 ≈ 50–80 concurrent engineers)Audited STAC-AI LANG6 results; MLPerf inference submissions
    AcceleratorsNVIDIA-validated, Blackwell-generation supportNVIDIA-first; AMD and Intel accelerator options as of mid-2026
    Best fitLarge vSphere estates staying with BroadcomKubernetes-first platform teams; VMware leavers

    Where VCF Private AI Services is strong

    Broadcom’s pitch is operational continuity, and it is genuinely persuasive for a vSphere-first shop. GPU-backed model workloads are managed like any other VCF workload — same lifecycle tooling, same admin skills, same DR posture your team already runs. The model store gives governance teams a curated, versioned catalog, which matters more than demos suggest once audit and the EU AI Act’s documentation requirements enter the picture. And the packaging change is the headline: because Private AI Services is included in VCF 9.0, the marginal software cost of a first AI workload on an existing estate is near zero. Your real spend is GPUs, power, and cooling — we have broken down that math in our on-prem GPU cluster cost guide.

    Who it fits: enterprises with 70%+ of workloads on vSphere, VMware-trained operations teams, and a signed multi-year VCF agreement. If that is you, the fastest governed pilot you can run this quarter is the one you have already licensed.

    Where OpenShift AI is strong

    Red Hat’s advantage is upstream gravity. vLLM has become the de facto open-source inference engine, and Red Hat employs a significant share of its maintainers following the Neural Magic acquisition — meaning optimizations land in OpenShift AI close to where they are invented. llm-d extends that with disaggregated, cache-aware serving across GPU pools, which is what large-scale multi-tenant inference actually requires. The audited STAC-AI results give risk-averse buyers — banks especially — something no internal vendor benchmark can: third-party verification on realistic workloads.

    Portability is the second argument. The same OpenShift AI platform runs on bare metal (no hypervisor layer on GPU nodes), on vSphere, or in public cloud. And for organizations exiting VMware over renewal economics, OpenShift Virtualization consolidates VMs, containers, and AI on one platform — we compared that path against the alternatives in our OpenShift Virtualization exit analysis. Who it fits: platform-engineering organizations already fluent in Kubernetes, and anyone whose three-year plan does not include Broadcom.

    The honest downsides

    VCF first. Private AI Services is young — several of its services shipped to general availability only this year, and the ecosystem around it is narrower than the Kubernetes AI ecosystem by an order of magnitude. It is also deeply coupled to NVIDIA; if you want accelerator optionality, this is not the stack. And the elephant: Broadcom’s per-core subscription model and licensing minimums have burned enough goodwill that “included at no extra cost” lands skeptically — the AI services are free, but the platform they require is not getting cheaper.

    OpenShift AI has its own tax. It demands real Kubernetes and MLOps skills — teams without them ship their first model months later than they planned. Total cost stacks OpenShift plus the OpenShift AI subscription, and list pricing varies enough by deal size that you should model both platforms per GPU node, not per socket. Red Hat’s AI portfolio — RHEL AI, OpenShift AI, AI Inference Server — also takes genuine effort to map onto a buying decision. And a caution that applies to both: every dollar here concentrates risk on NVIDIA supply, pricing, and the separately licensed NVIDIA AI Enterprise layer. Neither vendor shields you from that.

    What to do about it

    • VCF shops staying put: pilot Private AI Services on VCF 9.0 this quarter. It is already in your entitlement; the only new line items are GPUs and NVIDIA AI Enterprise.
    • Kubernetes-first or exiting VMware: choose OpenShift AI and fold the pilot into your migration business case — one platform decision instead of two.
    • Sizing rule of thumb: plan on one H100-class GPU per 50–80 concurrent code-assist users; halve that for RAG-heavy workloads, and validate with your own prompts before the purchase order goes out.
    • Check the buy-vs-API crossover first: below sustained utilization, hosted APIs still win on cost — run the numbers in our self-hosted LLM vs. API cost analysis before committing capital.
    • EU AI Act deadline: August 2026 enforcement makes model provenance, logging, and documentation table stakes. Score both platforms’ governance tooling against your compliance checklist, not the vendor demo.

    The repeatable line for your next steering meeting: the private AI platform decision is a platform-loyalty decision. Pick the stack your operations team can run at 2 a.m., because the models will change faster than the infrastructure underneath them.

    Frequently asked questions

    What is VMware Private AI Foundation with NVIDIA?

    It is the joint Broadcom–NVIDIA architecture for running generative AI on VMware Cloud Foundation with NVIDIA GPUs and NVIDIA AI Enterprise software. In VCF 9.0 it has evolved into Private AI Services — model store, model runtime, agent builder, vector database, and GPU monitoring — included in the VCF subscription by default.

    Is OpenShift AI better than VMware for private AI?

    Neither is universally better. OpenShift AI leads on open-source inference technology (vLLM, llm-d) and audited benchmarks; VCF Private AI Services leads on operational simplicity for existing vSphere estates and on packaging, since it is bundled with VCF 9.0. Choose based on whether your team’s core skill set is Kubernetes or vSphere.

    How many GPUs do I need to run a private LLM?

    As a planning heuristic, VMware’s internal benchmarking showed a single H100 supporting 50–80 concurrent engineers on code-assist inference. A quantized 7–13B model fits on one modern data-center GPU; 70B-class models generally need multiple GPUs or a distributed serving layer such as llm-d. Benchmark with your own workload before buying.

    Do both platforms require NVIDIA GPUs?

    VCF Private AI Services is an NVIDIA-specific stack. OpenShift AI is NVIDIA-first but also supports AMD and Intel accelerators as of mid-2026, which gives it more hardware optionality if GPU supply or pricing becomes a constraint.

    Does the EU AI Act require on-premises AI?

    No. The Act mandates governance, documentation, and risk controls — not where models run. On-premises deployment can simplify data-residency and confidentiality arguments, which is why enforcement beginning in August 2026 is accelerating private AI projects, but a well-governed cloud deployment can also comply.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Kubernetes Cost Optimization in 2026: Where Waste Hides

    Kubernetes Cost Optimization in 2026: Where Waste Hides

    Pull the utilization report for your busiest production cluster and brace yourself. Across the industry, average cluster CPU utilization sits near 10–13% of requested capacity — meaning most enterprises pay for roughly eight cores to use one. In practitioner surveys, around 70% name overprovisioning as the single biggest driver of Kubernetes overspend, and organizations without FinOps discipline waste an estimated 32–40% of their cloud bill versus 15–20% for mature teams. This guide walks through where that waste actually hides in 2026 — including the newest failure mode, four autoscalers silently fighting each other in the same cluster — and what AWS, Google Cloud, and Red Hat each give you to fix it.

    The waste numbers, quantified

    Before touching a single autoscaler, benchmark yourself. These are the thresholds I use when a VP asks whether their Kubernetes spend is normal.

    WastefulTypicalDisciplined
    CPU utilization vs. requestsUnder 15%15–35%40–60%
    Memory utilization vs. requestsUnder 25%25–50%50–70%
    Cloud spend wasted32–40%20–30%15–20%
    Spot / preemptible share of eligible workloads0%10–30%50%+

    The diagnostics are simple. If CPU utilization against requests is below 15%, your problem is requests, not node pricing — no reserved-instance negotiation will save you from an engineer who asked for 4 cores and uses 200 millicores. If you are between 15% and 35%, you have a right-sizing backlog worth real money. Above 40%, shift attention to purchase commitments and workload placement, because the easy waste is gone. And remember that compute is not the whole bill — data transfer stacks on top, which is why we benchmarked cloud egress fees across the big three separately.

    Why is overprovisioning so universal? Because the incentives run one way. The engineer who under-requests gets paged at 2 a.m. for OOMKills; the engineer who over-requests gets nothing but a quiet cluster. Requests get copied from a template written three years ago, load-tested never, and inflated after every incident. The takeaway for the meeting: overprovisioning is a management problem wearing a YAML costume. About 70% of practitioners agree it is the top overspend driver, and it will not fix itself.

    The 2026 failure mode: four autoscalers fighting each other

    The new waste pattern I see in 2026 audits is not missing automation — it is too much of it. A mature cluster now commonly runs four scaling systems at once: Cluster Autoscaler or Karpenter managing nodes, the Horizontal Pod Autoscaler scaling replica counts, and the Vertical Pod Autoscaler adjusting requests. Each is fine alone. Together, unconfigured, they fight.

    The classic loop: VPA raises a deployment’s CPU request, HPA — which scales on CPU utilization relative to that same request — suddenly sees utilization drop and scales replicas down, latency climbs, HPA scales back up, and the node autoscaler churns instances underneath the whole argument. Every cycle provisions capacity someone pays for. The symptoms are node counts that oscillate hourly without traffic changes, and consolidation that never converges.

    • Never point HPA and VPA at the same metric on the same workload — use VPA for request correction on steady services, HPA for genuinely elastic ones.
    • If you run Karpenter alongside a legacy Cluster Autoscaler node group, fence them to disjoint workloads or one will strand the other’s capacity.
    • Set consolidation policies deliberately — aggressive consolidation plus slow pod disruption budgets equals thrash.

    AWS: Karpenter and EKS Auto Mode

    Karpenter remains the strongest single cost lever on AWS. Instead of scaling fixed node groups, it provisions exactly the instances pending pods need, picks from the full EC2 catalog including Spot, and actively consolidates — repacking workloads onto fewer, cheaper nodes and terminating the rest. Since the v1 API (NodePool and EC2NodeClass replaced the old Provisioner objects), configuration is cleaner, and AWS has kept shipping meaningful speed improvements to scale-out and consolidation through 2025–26. Teams moving from Cluster Autoscaler to a well-tuned Karpenter setup routinely report double-digit percentage compute savings; the flexible-instance Spot story is where most of it comes from.

    EKS Auto Mode, introduced in late 2024, runs Karpenter as a managed component — AWS operates the controller, patches it, and handles node lifecycle. That is genuinely useful for teams that do not want to own autoscaler operations, and as of mid-2026 it has matured steadily, adding observability into scheduling decisions and zonal-failure handling. The honest downside: Auto Mode carries a management premium on top of EC2 pricing, and you give up some low-level tuning. If you have platform engineers who know Karpenter, self-managed is cheaper. If you do not, Auto Mode’s premium is usually smaller than the waste it removes. Where EKS sits against the other platforms overall is a bigger question — see our OpenShift vs. EKS vs. GKE comparison for that verdict.

    Google Cloud: Autopilot changes the unit of waste

    Google Cloud’s GKE Autopilot takes a different position: stop billing for nodes at all. You pay per pod for the CPU, memory, and ephemeral storage your pods request, in one-second increments, plus a cluster management fee of roughly $0.10 per hour — list pricing varies by region. That eliminates two entire waste categories, unused node headroom and bin-packing loss, because Google absorbs them.

    But notice what it does not eliminate: request inflation. Under Autopilot you pay for exactly what you request, so a service requesting 4 cores and using 200 millicores wastes money at 100% efficiency of billing. Autopilot converts the overprovisioning problem from an infrastructure problem into a pure right-sizing problem — which is progress, because right-sizing is measurable per team. Google’s per-vCPU Autopilot rates carry a premium over equivalent raw Compute Engine capacity, so well-utilized Standard-mode clusters (above roughly 50–60% utilization) often beat Autopilot on price. Below that, Autopilot usually wins. Spot Pods at up to around 60% off and flexible committed-use discounts sweeten it further for fault-tolerant and steady workloads respectively.

    Red Hat OpenShift: right-sizing inside the platform

    Red Hat’s angle is different again: OpenShift customers already pay for the platform, so cost tooling ships as part of the subscription rather than as a separate product. Red Hat Insights cost management aggregates spend across OpenShift clusters — on-prem and cloud — and its resource optimization service generates per-container right-sizing recommendations from observed usage. For enterprises running OpenShift on-prem or hybrid, that visibility matters more than any autoscaler, because on-prem waste hides in hardware refresh cycles instead of a monthly bill.

    The strengths: one cost view across hybrid estates, tight integration with cluster and machine autoscaling on supported clouds, and recommendations that account for OpenShift’s own overhead. The weaknesses are equally real — OpenShift subscription costs are themselves a significant line item, typically priced per core, so the platform raises your baseline even as its tooling trims usage. Red Hat fits organizations that have already committed to OpenShift for other reasons; nobody should buy it as a cost-optimization play. If that describes your shop, turning on cost management is free money you are currently leaving on the table.

    GPU workloads break bin-packing

    Every assumption above degrades when GPUs enter the cluster. Autoscalers optimize by treating CPU and memory as divisible commodities; GPUs are lumpy, expensive, and — without extra work — allocated whole. A pod that needs 20% of an accelerator strands the other 80%, and at current accelerator prices that stranded slice can cost more than an entire CPU node. Consolidation logic makes it worse: repacking a GPU workload means minutes of model reload, so disruption budgets block the very moves that save money elsewhere.

    The mitigations exist but require deliberate setup: fractional GPU sharing (time-slicing or MIG partitioning on supported NVIDIA hardware), separate node pools with their own scaling policies, and queue-based scheduling for batch inference and training. All three vendors are investing here — AWS via Karpenter GPU-aware provisioning, Google Cloud via Autopilot accelerator classes, Red Hat via OpenShift AI scheduling — but as of mid-2026 none of them makes fractional GPU economics automatic. Budget engineering time or budget stranded silicon.

    What to do about it

    Sequence matters. Right-size first, then automate, then commit — commitments made against inflated baselines lock waste in for one to three years.

    • Weeks 1–2: measure utilization against requests per namespace. Publish the league table internally. Shame is an underrated FinOps tool.
    • Month 1: right-size the top 20 offenders using VPA recommendations (recommendation mode, not auto mode, on HPA-scaled services).
    • Month 2: deploy or tune the node layer — Karpenter or Auto Mode on EKS, Autopilot evaluation on GKE, cost management plus autoscaling on OpenShift. Audit autoscaler interactions explicitly.
    • Month 3: move fault-tolerant workloads to Spot or Spot Pods, then buy commitments against the now-honest baseline.

    The takeaway for the meeting: a typical enterprise cluster runs near 10–13% CPU utilization, and getting to 40% is not heroic — it is three months of unglamorous work that usually returns 25–35% of the Kubernetes bill.

    Frequently asked questions

    How do I reduce Kubernetes costs quickly?

    Right-size resource requests first — that is where 70% of practitioners say the waste lives. Then enable node consolidation (Karpenter on AWS, Autopilot on GKE), move fault-tolerant workloads to Spot capacity, and only then buy committed-use discounts against the corrected baseline.

    Is Karpenter better than Cluster Autoscaler for cost savings?

    For most AWS workloads, yes. Karpenter’s instance-flexible provisioning, Spot integration, and active consolidation typically beat fixed node groups by a double-digit percentage. Cluster Autoscaler remains reasonable for small, stable clusters where churn is rare.

    Is GKE Autopilot cheaper than Standard mode?

    It depends on your utilization. Autopilot’s per-pod rates carry a premium over raw node capacity, so a Standard cluster running above roughly 50–60% utilization is usually cheaper. Below that — which is most clusters — Autopilot tends to win because you stop paying for idle headroom.

    What is a good CPU utilization target for Kubernetes?

    Aim for 40–60% of requested CPU actually used, measured over a week. Higher invites throttling and noisy-neighbor incidents; lower means you are funding idle cores. Memory targets run higher, around 50–70%, because memory exhaustion fails harder than CPU contention.

    Do I need a FinOps team to control Kubernetes spend?

    You need FinOps practices more than a FinOps org chart. The data is blunt: organizations with no cost discipline waste 32–40% of cloud spend versus 15–20% for mature teams. One engineer with showback dashboards and executive backing captures most of that gap.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Zscaler + Red Canary: SSE Vendor Moves Into the SOC

    Zscaler + Red Canary: SSE Vendor Moves Into the SOC

    On August 1, 2025 — the opening days of its fiscal 2026 — Zscaler closed its acquisition of Red Canary, a deal announced at roughly $675 million and closer to $692 million in total consideration once the filings did the math. On paper it looks like an SSE vendor buying a managed detection and response shop. In practice it is something more ambitious: an inline-proxy company betting it can build an AI-driven, largely autonomous SOC and stop being priced like a single-category security vendor.

    Nearly a year in, there is enough evidence to judge the bet. This brief covers what actually changed, whether Zscaler can credibly fight CrowdStrike, Palo Alto Networks, and Microsoft in security operations, what happens to Red Canary’s prized vendor neutrality — and what existing Red Canary MDR customers should renegotiate right now.

    What changed

    Red Canary now operates as a business unit inside Zscaler, keeping its brand but not its independence. Zscaler’s stated plan is to fuse three assets: Red Canary’s decade of MDR runbooks and its agentic AI investigation pipeline; the Zero Trust Exchange, which by the company’s own count inspects more than 500 billion transactions daily; and the Data Fabric for Security that came from the 2024 Avalor acquisition. The pitch is an agentic SOC — AI agents that triage, investigate, and increasingly respond to alerts with human analysts supervising rather than grinding through queues.

    Zscaler also now sells MDR under its own name, something it never did before. CEO Jay Chaudhry has framed the move as complementing rather than competing with MDR providers, but the product page tells its own story. The takeaway: Zscaler is no longer just the vendor that secures your traffic — it wants to be the vendor that runs your detection and response.

    Why it matters

    Security operations is where the consolidation war is being decided. CrowdStrike, Palo Alto Networks, and Microsoft have all spent the last three years pulling SIEM, XDR, and SOAR into single platforms, because whoever owns the SOC workflow owns the renewal conversation for everything else. Zscaler had a strong SSE franchise and no SOC story. That was a strategic dead end — the buyers I talk to increasingly want their zero-trust vendor and their operations vendor to at least share a data plane.

    Red Canary fixes the gap faster than Zscaler could have built it. The company was consistently rated among the strongest independent MDR providers, and its early, public work on agentic AI investigation was ahead of most of the market. What Zscaler bought is not just revenue — it is credibility with SOC teams, a population that historically viewed Zscaler as a networking purchase. Verdict: strategically sound, execution unproven.

    Can an inline-proxy vendor credibly run your SOC?

    Here is the honest problem. A SOC platform lives or dies on telemetry breadth, and Zscaler’s native telemetry is network-centric — web, SaaS, private app access, DNS. It sees traffic in extraordinary volume, but it does not natively own the endpoint, the identity tier, or the cloud workload the way its new rivals do. CrowdStrike starts from the endpoint, Microsoft from identity and productivity, Palo Alto from the firewall estate plus Cortex agents. Every one of them will argue their vantage point matters more than inline traffic.

    Zscaler’s counter is that Red Canary was built to be telemetry-agnostic — it ingests CrowdStrike, Microsoft Defender, SentinelOne, and others, and correlates across them. Pair that with the Data Fabric and Zscaler does not need to own every sensor, just the analytical layer above them. It is the same architectural argument we examined in our Zscaler vs Palo Alto SASE comparison: Zscaler wins when it positions as the neutral fabric, and struggles when it tries to out-platform the platforms. The SOC market will test that thesis harder than SASE ever did.

    The vendor-neutrality question

    Red Canary’s entire brand was built on independence — it graded EDR telemetry honestly because it sold none of its own. That stance is now structurally compromised: the parent company competes, at least partially, with the vendors whose telemetry Red Canary ingests. The expanded Zscaler–CrowdStrike partnership announced around the deal is a genuine mitigant, and dropping CrowdStrike support would be commercial suicide given how much of Red Canary’s customer base runs Falcon. Expect support to continue.

    But watch the roadmap, not the press releases. The predictable drift is that new agentic capabilities land first — and work best — when Zscaler telemetry and the Data Fabric are in the loop, while third-party integrations get maintenance-mode attention. Nothing malicious required; that is just how acquired platforms evolve. If you chose Red Canary specifically because it was Switzerland, that reason is expiring on a schedule nobody will announce.

    How the field lines up

    As of mid-2026, four credible platform paths exist for an AI-assisted SOC. The deeper feature-by-feature breakdown is in our AI SOC platforms comparison; the strategic shape is below.

    Zscaler + Red CanaryCrowdStrikePalo Alto NetworksMicrosoft
    SOC anchorMDR service + agentic investigation, Data Fabric analyticsFalcon platform, Next-Gen SIEM, Falcon Complete MDRCortex XSIAM “machine-led SOC” + Unit 42 servicesSentinel + Defender XDR, Security Copilot agents
    Native telemetry edgeInline traffic at massive scale; endpoint via third partiesEndpoint and identity depth; strong first-party sensorsFirewall estate plus Cortex endpoint agentsIdentity, email, productivity — the breach entry points
    Agentic AI maturityStrong — Red Canary shipped agentic triage earlyStrong — Charlotte AI expanding across workflowsStrong pitch; automation depth varies by data onboardingBroad agent catalog, uneven polish across it
    Best fitZscaler SSE shops; teams wanting MDR outcomes, not toolingEndpoint-first orgs consolidating onto FalconLarge SOCs ready to replace the SIEM outrightE5-heavy estates optimizing license economics
    Watch out forIntegration risk; neutrality drift; new to first-party SecOpsPremium pricing as module count growsMigration effort; services-heavy deploymentsMulti-cloud and non-Microsoft telemetry gaps

    The short version: Zscaler is the only one of the four entering the SOC from the service side rather than the tooling side. That is a real differentiator — many mid-enterprise teams want detection and response as an outcome, not another console — but it also means Zscaler’s SecOps revenue depends on people and AI agents performing, quarter after quarter, not just software shipping.

    What to do about it

    If you are an existing Red Canary MDR customer, act at renewal — not later. Three items belong in the negotiation. First, contractual protection for third-party telemetry: named support for your EDR of record (CrowdStrike, Defender, SentinelOne) for the full term, with service credits if parity slips. Second, pricing protection: acquired MDR services routinely get repackaged into platform bundles, and you want a cap on year-over-year increases before that happens. Third, data portability: confirm your detection history and tuned analytics can be exported if you leave.

    If you are a Zscaler SSE customer without Red Canary, the calculus is friendlier — bundling SSE with MDR from one vendor will likely be priced aggressively while Zscaler buys market share, and the integration genuinely shortens time-to-value if your traffic already flows through the Zero Trust Exchange. Pilot it against an incumbent quote. And if you are mid-evaluation for an AI SOC platform generally, do not let this deal rush you: run Zscaler + Red Canary head-to-head with at least one endpoint-anchored platform and score them on investigation quality per analyst-hour, not feature checklists.

    Frequently asked questions

    Why did Zscaler acquire Red Canary?

    Zscaler needed a security operations story to compete beyond SSE. Red Canary supplied a respected MDR business, ten years of detection runbooks, and early agentic AI investigation technology that Zscaler pairs with its Data Fabric for Security — accelerating a SOC roadmap that would have taken years to build internally.

    Will Red Canary still support CrowdStrike and Microsoft Defender telemetry?

    Yes, as of mid-2026 — Zscaler and CrowdStrike publicly expanded their partnership around the deal, and third-party EDR ingestion remains core to Red Canary’s offering. The open question is long-term roadmap parity, which is why customers should lock support commitments into renewal contracts.

    What is an agentic SOC?

    A security operations model where AI agents autonomously handle triage, enrichment, and investigation of alerts — escalating to human analysts for judgment calls and response authorization. The goal is cutting mean time to respond and analyst burnout, not eliminating the SOC team.

    Does Zscaler compete with CrowdStrike now?

    Partially. The companies remain integration partners and jointly serve many customers, but Zscaler’s MDR service now overlaps with CrowdStrike’s Falcon Complete, and both are chasing the same agentic SOC budget. Expect cooperation on telemetry and competition on services.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Google Closes Wiz: What CNAPP Consolidation Means

    Google Closes Wiz: What CNAPP Consolidation Means

    On March 11, 2026, Google closed its $32 billion acquisition of Wiz — the largest deal in Alphabet’s history, and the largest cybersecurity acquisition ever completed. It took a full year of regulatory review to get there. If your organization runs Wiz today, roughly one decision window just opened: the next 12–24 months will determine whether you renew with confidence, renegotiate hard, or migrate. This brief covers what actually changed, what the DOJ clearance does and does not guarantee, and how the credible alternatives — Palo Alto Networks, CrowdStrike, and Microsoft — stack up while you decide.

    What changed

    Google announced the Wiz deal in March 2025 at $32 billion in cash — after Wiz walked away from a reported $23 billion offer the previous summer. The DOJ’s antitrust review ran through late 2025, clearance came without conditions, and additional jurisdictions including the European Commission followed. The transaction closed March 11, 2026. Wiz now sits inside Google Cloud, alongside Mandiant ($5.4 billion, 2022) and Google Security Operations, giving Google the most expensive security portfolio ever assembled by a hyperscaler.

    Wiz was not a distressed asset. It went from founding in 2020 to hundreds of millions in ARR faster than any security company before it, on the strength of an agentless, graph-based approach to cloud risk that most of the Fortune 100 adopted. That is precisely why the deal matters: the independent CNAPP category just lost its flagship.

    Why it matters

    Consolidation shrinks your negotiating leverage. Three years ago a CNAPP shortlist had five or six independent vendors willing to discount aggressively to win the logo. Today the strongest independent is owned by a hyperscaler, Palo Alto Networks has folded Prisma Cloud into its broader Cortex platform, and Forrester has noted that late-stage security vendors now chase acquisitions rather than IPOs — meaning the remaining independents (Orca, Sysdig, Upwind and others) are themselves plausible acquisition targets. Every one of your alternatives is either a platform play or a future platform acquisition.

    The same consolidation logic is playing out across the security stack — we covered the parallel dynamic in SOC tooling in our Zscaler–Red Canary agentic SOC analysis. The pattern is consistent: acquirers buy best-of-breed tools, promise independence, then gradually tilt the roadmap toward their own platform. Sometimes the integration genuinely helps customers. It rarely helps their pricing.

    The takeaway for a VP: assume your Wiz renewal in 2027 is a negotiation with Google Cloud sales, not with a hungry startup. Plan leverage accordingly.

    The multicloud question — a promise, not a consent decree

    Here is the detail most coverage glossed over: the US clearance was reported as unconditional. Google has publicly and repeatedly committed that Wiz will remain multicloud — continuing to protect workloads on AWS, Azure, and Oracle Cloud — and the DOJ scrutinized exactly that question during its review. But a public commitment is not a binding remedy. There is no consent decree compelling Google to maintain feature parity for AWS-hosted workloads in 2028.

    To be fair to Google Cloud, its commercial incentives mostly point the right way. The majority of Wiz revenue comes from customers whose primary cloud is AWS or Azure; gutting multicloud support would destroy the asset it just paid $32 billion for. Google also has a defensible track record here — Mandiant still serves non-Google environments four years after acquisition. The realistic risk is not a shutdown. It is drift: new capabilities landing on Google Cloud first, deeper Security Command Center integration becoming the default posture, and AWS-specific coverage moving at a slower cadence. Watch release notes, not press releases.

    Where the alternatives stand

    Wiz (Google Cloud)Palo Alto Cortex CloudCrowdStrike Falcon Cloud SecurityMicrosoft Defender for Cloud
    Core approachAgentless graph-based posture, expanding runtime (Wiz Defend)CNAPP merged into Cortex SOC platformRuntime-first, single agent shared with EDR, plus agentless postureNative Azure CSPM/CWPP, bundled with licensing
    Strongest fitMulticloud estates that prioritized fast visibilityEnterprises standardized on Palo Alto network + SOC stackRuntime threat detection where the Falcon agent is already deployedAzure-first shops with E5 / Enterprise Agreements
    Watch out forRoadmap gravity toward Google Cloud; renewal leverage shifts to Google salesPrisma-to-Cortex migration mechanics; platform-bundle pricing pressurePosture/graph depth trails Wiz; strongest value requires the wider Falcon platformAWS/GCP coverage thinner than Azure; alert quality varies by plan tier

    Palo Alto Networks is the most direct beneficiary on paper — it has spent two years arguing for platform consolidation, and its move of Prisma Cloud into Cortex Cloud puts CNAPP, SIEM-successor, and SOC automation on one data plane. The honest caveat: existing Prisma Cloud customers describe the transition as a real migration, not a rebrand, and Palo Alto’s bundle-heavy discounting rewards customers who commit broadly. We saw the same platform-versus-point-product tension in our Zscaler vs Palo Alto SASE comparison — the pattern repeats in cloud security.

    CrowdStrike comes at CNAPP from runtime. Falcon Cloud Security uses the same agent an enterprise already runs for EDR, which makes container and workload protection close to free operationally, and Frost & Sullivan again ranked it a CNAPP leader in 2026. Its graph-style posture analytics are improving but still trail Wiz’s; CrowdStrike’s economics also work best when you buy the broader Falcon platform. Microsoft Defender for Cloud is the pragmatic default for Azure-first organizations — foundational CSPM effectively rides along with an Enterprise Agreement — but its AWS and GCP coverage remains noticeably thinner than its Azure depth, and Microsoft is a hyperscaler with the same conflict-of-interest profile people worry about with Google.

    The 12–24 month Wiz customer checklist

    • Contract terms (now): at renewal, push for multi-year price protection, a cap on uplift (5–7% is achievable in this market), and portability clauses — data export formats and reasonable termination assistance.
    • Roadmap independence (quarterly): track whether AWS and Azure connectors get new capabilities within the same quarter as Google Cloud. Two consecutive quarters of lag is your early-warning signal.
    • Account team continuity (6 months): if your Wiz account team is absorbed into Google Cloud field sales and your CNAPP renewal starts arriving bundled with GCP commit conversations, treat that as a structural change in the relationship.
    • Integration posture (12 months): confirm the third-party integrations you depend on — Splunk, ServiceNow, CrowdStrike, Microsoft Sentinel — remain first-class, not merely “supported.”
    • Benchmark an alternative (by month 18): run a scoped proof of concept with at least one of Cortex Cloud, Falcon Cloud Security, or Defender for Cloud before your renewal window, even if you intend to stay. A live alternative is worth 15–20% at the table.

    What to do about it

    If you are a happy Wiz customer on GCP, or genuinely multicloud with a Google tilt — stay. The product will likely get better for you, and Security Command Center integration is real upside. If you are AWS- or Azure-dominant, staying is still defensible, but only with the contract protections above; do not sign a three-year renewal in 2026 without them. If you were mid-evaluation when the deal closed, widen the shortlist: Palo Alto Networks for consolidated SOC platforms, CrowdStrike for runtime-first estates with Falcon already deployed, Microsoft for Azure-centric licensing economics, and the remaining independents if credible neutrality is a hard requirement — priced with the knowledge that they may be acquired too.

    The one-line version for your next steering meeting: Wiz under Google is not a risk event, it is a leverage event — and leverage is something you rebuild deliberately, starting at the next renewal.

    Frequently asked questions

    Is the Google acquisition of Wiz complete?

    Yes. The deal closed on March 11, 2026, after clearing the DOJ’s antitrust review in late 2025 and approvals in additional jurisdictions. At $32 billion in cash, it is the largest acquisition Alphabet has ever made.

    Will Wiz still support AWS and Azure after the Google acquisition?

    Google has committed publicly that Wiz remains multicloud, and most Wiz revenue depends on AWS- and Azure-hosted customers, so the incentive to keep that promise is strong. But the commitment is not a binding regulatory condition — customers should secure contractual protections and monitor whether non-GCP features keep pace.

    What are the best alternatives to Wiz in 2026?

    The three most credible enterprise alternatives are Palo Alto Networks Cortex Cloud (formerly Prisma Cloud), CrowdStrike Falcon Cloud Security, and Microsoft Defender for Cloud. Independent options such as Orca Security and Sysdig remain viable, particularly where vendor neutrality matters.

    Should Wiz customers switch vendors because of the acquisition?

    Not automatically. There is no immediate product risk, and Google’s Mandiant track record suggests continuity. The prudent move is to tighten renewal terms, track roadmap parity across clouds for 12–24 months, and keep one benchmarked alternative ready before the next renewal.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • AWS European Sovereign Cloud: Who Actually Needs It

    AWS European Sovereign Cloud: Who Actually Needs It

    On January 15, 2026, AWS switched on its European Sovereign Cloud — a physically and logically separate cloud with its first region in Brandenburg, Germany, more than 90 services at launch, and a commitment of over €7.8 billion in German investment behind it. Every operator with access to the environment is an EU resident. Identity, billing, and usage metadata never leave the EU. Sovereign Local Zones are already slated for Belgium, the Netherlands, and Portugal. This is the most serious sovereignty play a US hyperscaler has made, and I have sat through enough procurement reviews to know the question every architecture board will now ask: do we actually need this?

    Most coverage so far has been press-release paraphrase. This brief does the harder part — separating the controls that are genuinely new from the legal question no launch event resolved, comparing the Microsoft and Google Cloud sovereign approaches, and giving you a workload-level framework for when the premium is justified.

    What changed

    Until now, “sovereign” options from US hyperscalers were mostly policy wrappers: data residency commitments, EU support staff on the front line, contractual promises layered over shared global infrastructure. The AWS European Sovereign Cloud is structurally different. It is a separate cloud — its own regions, its own control plane, its own identity system, its own billing stack — operated by EU-resident personnel under an EU-based corporate structure. Even the metadata that normally flows to global systems, the stuff residency contracts quietly exclude, stays inside the EU.

    AWS also announced continuity assurances, including source code and technical documentation held in escrow within the EU so the environment can keep running under adverse scenarios. That escrow provision is the tell. AWS is trying to answer a question customers started asking loudly in 2025: what happens to our infrastructure if transatlantic relations deteriorate? A vendor building for that scenario is a genuine shift in posture.

    Why it matters

    The money says this is not a niche. European sovereign cloud spending grew roughly 83% year over year from a 2025 base near €6.9 billion, and worldwide sovereign cloud spending is forecast to reach $80 billion in 2026. Meanwhile US providers still hold more than 70% of the EU cloud market against roughly 15% for European providers. That tension — European regulatory pressure meeting entrenched US hyperscaler dependence — is exactly the gap this product is built to occupy.

    For EU public sector bodies and regulated industries, a fully featured sovereign environment from the market leader changes the default. Before January, choosing sovereignty usually meant choosing a smaller European provider and accepting a thinner service catalog. Now the trade-off is narrower. And for US multinationals with EU subsidiaries, the calculus runs the other way — a sovereign region is suddenly a credible answer when an EU regulator or customer asks where the workload actually lives and who can touch it. If your organization is also weighing whether some of these workloads belong in the cloud at all, our cloud repatriation analysis covers that adjacent decision.

    What AWS actually built — and what it didn’t solve

    Credit where due. The genuine controls: physical and logical separation from the global AWS partition, EU-only operations staffing, sovereign IAM so identity data never transits US systems, EU-resident billing and metering metadata, and the EU escrow arrangement. Launching with 90+ services means the catalog covers most of what a typical enterprise stack needs on day one — a real differentiator against sovereign offerings that launch with a dozen services and a roadmap.

    Here is what the launch did not solve: the operating entity is still an Amazon subsidiary. The unresolved question is whether US legal process — the CLOUD Act in particular — can compel a US parent to produce data held by an EU entity it ultimately owns. AWS has structured the ESC to make that maximally difficult, both technically and legally, and its position is that the controls put customer data beyond unilateral reach. But no court has tested this structure. Lawyers I trust call it a strong mitigation, not an elimination, of jurisdictional risk. If your threat model requires that a US court order be legally impossible rather than practically frustrated, an American-owned entity cannot get you there — only an EU-owned operator can. That is the honest boundary of this product, and buyers should price it in.

    How Microsoft and Google Cloud compare

    The three hyperscalers have taken visibly different routes to the same demand. Microsoft, since its June 2025 announcements, offers a Sovereign Public Cloud across its existing European regions — data under European law, European personnel controlling operations and access, customer-held encryption keys, plus tooling like Data Guardian and External Key Management — and a Sovereign Private Cloud built on Azure Local, including Microsoft 365 Local for productivity workloads on-premises. Google Cloud leans hardest on partner operation: its French Trusted Cloud runs through S3NS, the Thales joint venture, and in 2025–26 it added Google Cloud Dedicated for partner-operated regional deployments and a fully air-gapped option for disconnected environments.

    AWS European Sovereign CloudMicrosoft Sovereign CloudGoogle Cloud sovereign options
    ModelSeparate cloud partition, EU-operated, launched Jan 2026Sovereign controls layered on existing EU regions, plus private cloud on Azure LocalPartner-operated (S3NS/Thales), Dedicated, and air-gapped deployments
    Ownership of operatorAmazon subsidiary, EU-based structure, EU-resident staffMicrosoft, with European personnel controlsVaries — partner entities can be EU-owned
    Service breadth90+ services at launchBroad — spans Azure, M365, Power PlatformNarrower in partner and air-gapped models
    Strongest fitRegulated enterprises wanting full AWS depth with maximum isolationM365-centric estates needing sovereignty across productivity and cloudBuyers who require an EU-owned operating entity
    Honest weaknessUS parent ownership; untested against US legal processShared-region model is a weaker isolation story than a separate partitionFragmented catalog; partner model adds operational seams

    The verdict is workload-shaped. AWS now has the deepest isolated-but-full-featured offering. Microsoft has the broadest sovereignty story across productivity plus cloud — nobody else covers the M365 estate. Google Cloud’s partner structure is the only hyperscaler answer for buyers whose lawyers insist on an EU-owned operator, at the cost of catalog depth and an extra party in every escalation.

    Which workloads justify the premium

    Sovereign environments cost more — expect a meaningful premium over standard regions (list pricing varies; model it per workload) plus migration and dual-operating overhead. My rule of thumb after two decades of these reviews: sort workloads into three buckets.

    • Mandated: workloads where a regulator, national security framework, or contract explicitly requires sovereign operation — government, defense-adjacent, some health and critical infrastructure. The premium is the cost of doing business. Move these first.
    • Defensible: regulated-industry workloads (banking, insurance, pharma) where sovereignty reduces audit friction and future-proofs against tightening EU rules. Justify case by case — often the win is faster approvals, not compliance necessity.
    • Everything else: if no regulator, customer, or credible geopolitical scenario demands it, standard EU regions with residency controls remain the right answer. Paying the sovereign premium here is theater.

    One more filter: AI workloads deserve special attention, because training data and model governance obligations increasingly carry their own residency strings — our enterprise AI governance guide maps those requirements in detail.

    What to do about it

    Three actions this quarter. First, classify: run the three-bucket exercise above against your EU workload inventory before a vendor account team does it for you. Second, interrogate: ask each provider the CLOUD Act question in writing and compare the answers — the shape of the hedging is itself useful data. Third, price: get sovereign-region quotes for your mandated bucket now, because launch-window commercial flexibility is real and it fades. The takeaway for the meeting: sovereignty is now a workload attribute, not a vendor religion — buy it where it’s mandated, defend it where it’s defensible, and skip it everywhere else.

    Frequently asked questions

    Is the AWS European Sovereign Cloud subject to the US CLOUD Act?

    Unresolved. The operator is EU-based with EU-resident staff and strong technical controls, but it remains an Amazon subsidiary, and no court has tested whether US legal process can reach data held under this structure. Treat it as strong mitigation, not legal immunity.

    How is the AWS European Sovereign Cloud different from regular AWS EU regions?

    Standard EU regions keep customer data in-region but rely on global control planes, identity systems, and support. The Sovereign Cloud is a separate partition with its own EU-resident operations, sovereign IAM, and EU-held billing and usage metadata.

    How much more does a sovereign cloud cost?

    Expect a premium over standard regions — exact list pricing varies by service and commitment as of mid-2026. Budget for migration and operational duplication too, which often exceed the raw infrastructure delta.

    Who actually needs a sovereign cloud?

    Organizations under explicit regulatory or contractual sovereignty mandates — government, defense-adjacent, parts of health, finance, and critical infrastructure — plus multinationals whose EU customers demand it. For most other workloads, standard EU regions with residency controls are sufficient.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • One AI Governance Stack: EU AI Act, ISO 42001, NIST RMF

    One AI Governance Stack: EU AI Act, ISO 42001, NIST RMF

    August 2, 2026 is the date to circle. That is when the European Commission’s enforcement powers over general-purpose AI obligations kick in — obligations that technically applied back in August 2025 but had no enforcement teeth behind them. Penalties under the EU AI Act run up to €15 million or 3% of global turnover for most violations, and up to €35 million or 7% for prohibited practices. Meanwhile ISO 42001 has quietly become a procurement checkbox, and NIST AI RMF language is showing up in US federal contract clauses. Three frameworks, one budget cycle.

    The mistake I keep seeing: enterprises standing up three separate compliance programs, each with its own spreadsheet, its own owner, and its own steering committee. Run one governance stack instead — anchored on a single AI system register — and let each framework consume the evidence it needs. This guide covers the sequencing, the realistic timeline, and how your model vendors’ compliance paperwork feeds your evidence file.

    The 2026 deadline math

    The AI Act’s dates have moved, and most summaries you read in 2025 are now wrong. Prohibited practices and the AI literacy obligation have applied since February 2025. GPAI model obligations applied in August 2025. What changes on August 2, 2026 is enforcement: the Commission can start levying fines, and the Act’s transparency rules take effect. The relief valve is on the high-risk side — under the timeline amendments adopted in 2026, obligations for Annex III high-risk systems (the use-case-based tier: hiring, credit, critical infrastructure) were postponed to December 2027.

    Translation for a VP: if you deploy AI in the EU, your exposure through 2026 is mostly transparency, AI literacy, and — if you build or meaningfully fine-tune models — GPAI documentation. The high-risk conformity work got a reprieve, but it did not go away. Starting that work in early 2027 for a December 2027 date is how programs fail. The takeaway: treat August 2026 as the enforcement start line, not the finish line.

    Why one stack beats three programs

    The three frameworks overlap far more than their acronyms suggest. All three demand the same core artifacts: an inventory of AI systems, a risk classification per system, documented human oversight, incident handling, and lifecycle monitoring. Where they differ is posture — NIST AI RMF is voluntary guidance, ISO 42001 is a certifiable management system, and the EU AI Act is law with fines. Build the artifacts once, map them three ways.

    NIST AI RMFISO/IEC 42001EU AI Act
    What it isVoluntary risk framework (Govern, Map, Measure, Manage)Certifiable AI management system standardBinding regulation with penalties
    Who demands itUS federal contracts, boards, insurersEnterprise procurement, RFPsAny org placing AI on the EU market or using it there
    Core artifactRisk profile per systemAIMS documentation + audit evidenceTechnical documentation, conformity assessment, registration
    ProofSelf-attestationThird-party certificateCE marking / registration, regulator scrutiny
    Timeline pressureContract-driven, nowProcurement-driven, rising through 2026Aug 2026 enforcement; Dec 2027 for Annex III high-risk

    Run three separate programs and you will pay three times for the same inventory work, and the documents will drift out of sync — which is exactly what an auditor or regulator will find. One stack, one register, three reporting views. That sentence is the whole strategy.

    The AI system register is the anchor

    Every framework conversation eventually collapses into one question: do you actually know what AI is running in your enterprise? The register is the answer, and it must be a living system — not a quarterly spreadsheet. Minimum fields per entry:

    • System name, business owner, and technical owner — a name, not a team alias
    • Underlying model and provider (including version), plus deployment path — API, cloud marketplace, or self-hosted
    • Risk classification under each framework: EU AI Act tier, ISO 42001 impact assessment, NIST risk profile
    • Data touched, human oversight mechanism, and links to the vendor’s compliance documentation
    • Review date and incident log pointer

    Two failure modes to design against. First, shadow AI — the register only covers what you can see, so pair it with discovery controls; we covered the policy side in our shadow AI risk analysis. Second, agents — once systems act autonomously with their own credentials, the register needs identity fields too, which is the subject of our agent identity governance guide. A register that misses either is a compliance prop, not a control.

    Sequencing: NIST first, ISO second, EU AI Act layer

    For a moderately complex organization — a few dozen AI systems, one or two jurisdictions — plan on 8–12 months end to end. The order matters:

    PhaseTypical durationExit criteria
    1NIST AI RMF adoption: register built, risk profiles, governance roles named2–4 monthsEvery production AI system has an owner and a risk profile
    2ISO 42001: formalize the RMF work into an auditable AIMS, then certify4–6 monthsStage 2 audit passed; certificate in hand
    3EU AI Act layer: map register entries to Act tiers, close gaps for in-scope systems2–3 months (only if EU-exposed)Transparency and documentation obligations evidenced per system

    Why this order: NIST RMF is the cheapest place to make mistakes — it is voluntary, so you can iterate on the register and roles without an auditor watching. ISO 42001 then certifies discipline you already practice rather than discipline you invented for the audit. The EU AI Act layer goes last because it is a mapping exercise once the first two exist. Organizations that start with the Act tend to lawyer the problem instead of engineering it. If you run above roughly 50 AI systems or heavily regulated workloads, add 3–4 months and budget for external help on the conformity side.

    What your model vendors hand you

    You inherit a large chunk of your evidence file from your model providers — if you know where to look. As of mid-2026, all four major providers hold ISO/IEC 42001 certification, which materially shortens your own vendor-risk workload. What differs is the depth and shape of what each hands you.

    Microsoft has the broadest certified surface: ISO 42001 coverage spans GitHub Copilot, Microsoft 365 Copilot, Security Copilot, and the Foundry platform, and Purview Compliance Manager ships assessment templates that map controls to the EU AI Act and ISO 42001. If you are an M365/Azure shop, this is the shortest path to a populated evidence file. The honest caveat: Microsoft’s governance tooling maps best to Microsoft’s own estate — bring third-party or self-hosted models and you are back to manual mapping.

    Google certified its AI management system across Google Cloud, Workspace, and the Gemini app, and its model cards for Gemini remain among the more rigorous public documentation in the industry. Strong fit for GCP-first shops and for teams that want per-model technical detail. The weakness is coherence — evidence lives across several consoles and documentation sites, so assembling a per-system package takes more assembly work than it should.

    Anthropic was among the first frontier labs to achieve accredited ISO 42001 certification (announced January 2025), and its system cards for Claude models are detailed enough to lift directly into EU AI Act technical documentation. Fit: enterprises that want a defensible paper trail on model behavior and safety testing. Limitation: as a model provider rather than a hyperscaler, Anthropic hands you model-level evidence — platform-level controls come from wherever you deploy, such as Bedrock or Vertex.

    OpenAI maintains ISO 42001 coverage across its consumer and business products and publishes system cards per frontier model. The documentation is solid; the operational challenge is churn. Product names, model versions, and data-handling terms have shifted quickly, which means evidence you filed six months ago may reference a product that no longer exists under that name. Assign someone to re-verify OpenAI-derived evidence quarterly.

    Rule of thumb: vendor certificates cover the model and platform layer, never your use of it. A certified model in an uncontrolled workflow is still your liability. And if data-sovereignty pressure pushes you toward self-hosting, the governance calculus changes again — see our comparison of private AI on VMware versus OpenShift AI.

    The people layer: AI literacy evidence

    The most-ignored EU AI Act obligation is Article 4: AI literacy, in force since February 2025. Regulators will ask how you ensured staff using AI systems understand them — and “we sent an email” is not evidence. This is where training platforms earn a place in the governance stack. KnowBe4, best known for security awareness training, fits here: its compliance training library and human risk management platform give you assignable AI-use modules with completion tracking, which is precisely the auditable artifact Article 4 and ISO 42001’s competence clauses want. Strengths: enterprise-scale rollout and LMS-grade evidence trails your auditor already understands. Limits: KnowBe4 is a human-layer control, not a governance platform — it will not build your register, classify your systems, or manage conformity. Use it for the literacy evidence line, and do not let a completed training campaign masquerade as a governance program.

    Frequently asked questions

    Do I need ISO 42001 if I already follow NIST AI RMF?

    Legally, no — but procurement increasingly says yes. NIST RMF is self-attested; ISO 42001 is a third-party certificate a customer can verify. If you sell to enterprises or governments, expect the certificate to become table stakes in RFPs through 2026–27. The good news: an honest RMF implementation gets you most of the way there.

    Does the EU AI Act apply to US companies?

    Yes, if you place AI systems on the EU market or their outputs are used in the EU. Like GDPR, it is extraterritorial. A US SaaS product with EU customers is in scope even with zero EU infrastructure.

    How long does ISO 42001 certification take?

    Plan 4–6 months from a working governance baseline to a passed Stage 2 audit, plus lead time to book an accredited certification body — audit capacity has been tight as demand rose through 2026. Starting from nothing, treat the full journey as 8–12 months.

    What happens on August 2, 2026?

    The Commission’s enforcement powers over GPAI obligations begin and the Act’s transparency rules take effect. Most Annex III high-risk system obligations were pushed to December 2027 by the 2026 timeline amendments — a reprieve for deployers, not a cancellation.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • NVIDIA AI Factories: What the Reference Designs Signal

    NVIDIA AI Factories: What the Reference Designs Signal

    NVIDIA now publishes Enterprise Reference Architectures for three distinct classes of on-prem “AI factory” — and if you are budgeting GPU infrastructure for 2026–27, those three documents will shape your shortlist more than any RFP you write. The tiers run from RTX PRO servers for departmental inference, through HGX-based systems in the mid-range, up to rack-scale GB300 NVL72 machines with 72 Blackwell Ultra GPUs and 36 Grace CPUs per rack, stitched together with Spectrum-X networking and BlueField DPUs. Validated designs starting at roughly four-node clusters mean the entry point is no longer a hyperscaler-only conversation.

    “AI factory” is a marketing term. Underneath it is a real architectural decision with a ten-year procurement shadow. This brief decodes the tiers into a sizing decision tree, explains why the full-stack packaging should remind you of converged infrastructure circa 2014, and covers where Broadcom/VMware and Red Hat fit — because the software layer you pick determines how locked in the hardware layer leaves you.

    What changed

    NVIDIA’s Enterprise Reference Architectures (ERAs) have matured from loose sizing guides into full validated designs, and as of mid-2026 they define three named AI factory configurations: the RTX PRO AI Factory built on RTX PRO servers, the HGX AI Factory built on NVIDIA-Certified HGX systems, and the NVL72 AI Factory built on rack-scale GB300 NVL72 platforms. Each document prescribes the compute nodes, the Spectrum-X Ethernet fabric, BlueField DPU placement, storage partner options, and the NVIDIA AI Enterprise software stack on top. OEMs — Dell, HPE, Lenovo, Supermicro, Cisco — ship systems certified against these designs rather than inventing their own.

    The practical shift: an enterprise no longer buys GPUs; it buys a validated cluster with the network and software pre-decided. NVIDIA claims the GB300 NVL72 delivers up to 50× the AI factory output of Hopper-generation systems and 30× faster real-time inference on trillion-parameter models. Treat vendor multipliers with the usual skepticism — but the direction is not in dispute. Rack-scale is the new unit of purchase at the top end.

    The three tiers, decoded

    RTX PRO AI FactoryHGX AI FactoryNVL72 AI Factory
    Compute unit2U RTX PRO servers, air-cooledNVIDIA-Certified HGX nodes (8 GPUs/node class)GB300 NVL72 rack: 72 Blackwell Ultra GPUs, 36 Grace CPUs
    Entry scale~4-node clusters upwardSmall multi-node clusters to hundreds of GPUsOne rack minimum; multi-rack pods
    Primary workloadsDepartmental inference, RAG, VDI-adjacent AI, fine-tuning small modelsSerious fine-tuning, mid-size training, high-throughput inferenceFoundation-model training, trillion-parameter reasoning, agentic pipelines
    Facilities impactStandard racks and powerHigh-density racks; liquid cooling increasingly assumedLiquid cooling and 100kW+ rack power, full stop
    Who it fitsMost enterprises starting private AIEnterprises with committed AI product roadmapsModel builders, sovereign AI, GPU service providers

    The honest read on each tier. The RTX PRO tier is the one most IT shops should study first — it runs on facilities you already have, and for inference-heavy private AI (which is what most enterprise AI actually is) it is usually enough. The HGX tier is the default answer when data science teams demand training capacity; it is also where cost overruns live, because the network and storage bill surprises people. The NVL72 tier is genuinely impressive engineering, but if you have to ask whether you need it, you do not. Run the arithmetic in our on-prem GPU cluster cost guide before any of these conversations.

    Why it matters

    We have seen this movie. A decade ago, converged and hyperconverged infrastructure took the server-network-storage decision away from component buyers and sold it as one validated SKU. It worked — deployment risk fell, time-to-production fell — and vendor optionality fell with it. NVIDIA’s ERAs do the same for AI: compute, Spectrum-X networking, BlueField DPUs, and NVIDIA AI Enterprise software arrive as a package, and every layer you accept is a layer you will not competitively bid later.

    The networking clause is the one to read twice. The reference designs standardize on Spectrum-X Ethernet, which is excellent — and which quietly displaces the Arista, Cisco, or Juniper fabric your network team would otherwise have specified. Same with BlueField DPUs for east-west security and storage offload. None of this is bad engineering; all of it is deliberate account expansion. The takeaway for the meeting: validated designs buy you speed today at the price of negotiating leverage in 2029.

    Where Broadcom/VMware fits

    Broadcom’s answer is VMware Private AI Foundation with NVIDIA, which layers vGPU-partitioned NVIDIA AI Enterprise onto VMware Cloud Foundation. Its strength is real: if you are already a VCF shop, your operations team keeps the tooling it knows — vCenter, DRS, snapshots, the whole virtualization discipline — while data scientists get self-service GPU workstations and model runtimes. GPU sharing via vGPU is the underrated feature; departmental inference rarely saturates a full GPU, and virtualization claws that waste back.

    The weaknesses are equally real. You are stacking two aggressive licensing regimes — Broadcom’s post-acquisition VCF subscription pricing plus NVIDIA AI Enterprise — on the same cluster, and Broadcom’s pricing changes since 2024 have made renewal math painful for mid-size shops. It fits large VMware-committed enterprises running mixed workloads; it fits poorly if you are trying to exit VCF or if your AI platform will be container-native from day one. Our VMware vs. OpenShift AI comparison works through that decision in detail.

    Where Red Hat fits

    Red Hat’s play is OpenShift AI: a Kubernetes-native MLOps platform where the NVIDIA GPU Operator manages device lifecycle cluster-wide and NVIDIA AI Enterprise is a certified overlay rather than the operating model. Its strength is portability — the same OpenShift AI stack runs on bare metal, on vSphere, in public cloud, or at the edge, which makes it the natural hedge against exactly the lock-in the ERAs encourage. Model serving, pipelines, and open source model support (including Red Hat’s vLLM-based inference work) are first-class, not bolted on.

    The trade-off is operational: OpenShift assumes platform-engineering maturity. If your organization does not already run Kubernetes competently, standing up OpenShift AI alongside new GPU hardware is two transformations at once, and the failure mode is a very expensive science project. It fits container-fluent enterprises and regulated shops that need hybrid portability; it does not fit teams whose operational center of gravity is still the vSphere client.

    What to do about it

    A sizing decision tree you can defend in a budget meeting:

    • Workload is inference, RAG, or fine-tunes under ~70B parameters: start at the RTX PRO tier, four to eight nodes. Do not let anyone sell you an HGX cluster for a chatbot.
    • Sustained training or GPU utilization forecast above ~60% around the clock: the HGX tier beats cloud economics over a 3-year horizon; below that threshold, rent first and buy after you have utilization data.
    • Training foundation models or serving trillion-parameter reasoning at scale: NVL72 territory — and a facilities project before it is an IT project. Budget the liquid cooling retrofit honestly.
    • Negotiate the fabric separately. Accept the validated compute design, but make Spectrum-X win the network on price against at least one alternative bid, even if you expect it to win.
    • Pick the software layer for your team, not the diagram. VMware-committed and virtualization-strong: Private AI Foundation. Kubernetes-strong and portability-minded: OpenShift AI. Neither: fix that before buying racks.

    The one-line version for your VP: buy the smallest validated tier that covers the next 18 months, keep the network and software layers contestable, and let utilization data — not reference architecture diagrams — justify the next tier up.

    Frequently asked questions

    What is an NVIDIA AI factory?

    It is NVIDIA’s term for a data center built as a production line for AI — GPU compute, high-speed networking, storage, and software packaged to turn data into models and tokens. In enterprise practice it means a cluster built to one of NVIDIA’s Enterprise Reference Architectures rather than a hand-assembled GPU farm.

    What is the difference between GB300 NVL72 and HGX systems?

    HGX systems are conventional servers with up to eight GPUs each, clustered over a network. GB300 NVL72 is a single liquid-cooled rack where 72 Blackwell Ultra GPUs and 36 Grace CPUs behave as one NVLink-connected accelerator — bought, powered, and operated as a rack-scale unit.

    Do I need NVIDIA AI Enterprise to build an AI factory?

    The reference architectures assume it, and both VMware Private AI Foundation and Red Hat OpenShift AI integrate it. You can run open source stacks on the same hardware, but you give up the certified support matrix — a trade most regulated enterprises decline.

    Is on-prem AI infrastructure cheaper than cloud GPUs?

    Only above a utilization threshold — as a rule of thumb, sustained utilization north of roughly 60% over three years favors owning, and spiky or exploratory workloads favor renting. Data gravity, sovereignty rules, and per-token inference volume shift the math case by case.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Bedrock vs. Vertex AI vs. Azure AI Foundry in 2026

    Bedrock vs. Vertex AI vs. Azure AI Foundry in 2026

    By mid-2026, the three hyperscaler AI platforms have copied each other’s homework. Managed model catalogs, agent frameworks, guardrails, RAG tooling, provisioned throughput — pull a feature slide from AWS Bedrock, Google Vertex AI, or Microsoft’s Azure AI Foundry and you could swap the logo without anyone noticing. The checklist era is over.

    What still separates them is much harder to copy: which frontier models each cloud is allowed to serve you, and which compliance authorizations each has actually cleared. Those two factors — plus a pair of billing traps that surprise teams after the contract is signed — should drive the decision in 2026. This brief covers all of it.

    The short version

    If your enterprise runs regulated workloads that require FedRAMP High or DoD impact levels, the choice is Bedrock or Azure OpenAI — Vertex is disqualified until its authorization goes GA, regardless of how good Gemini is. If you need OpenAI’s frontier models under an enterprise agreement, Azure AI Foundry is the only cloud that has them. If model breadth is your priority and compliance isn’t a gate, Vertex AI’s Model Garden is the deepest catalog of the three. And if you want a hedge, note that Anthropic’s Claude family is the one frontier lineup available on all three clouds.

    AWS BedrockGoogle Vertex AIAzure AI Foundry
    Anchor modelsClaude, Amazon NovaGeminiOpenAI frontier models
    Also servesLlama, Mistral, Cohere, open-weight GPT OSSClaude, Gemma, Llama, Mistral — 50+ in Model GardenClaude, Llama, Mistral, Grok, DeepSeek
    OpenAI frontier modelsNo (open-weight variants only)NoYes — exclusive among the clouds
    FedRAMP HighYes, incl. GovCloud IL-4/5Not yet GA as of mid-2026Yes (Azure OpenAI)
    Billing watch-outProvisioned throughput commitmentsThinking tokens billed as output; long-context rate step-upCapacity quotas on frontier models
    Best fitRegulated workloads, AWS-committed shopsModel breadth, Google-stack shopsOpenAI dependency, Microsoft-committed shops

    What changed

    Three shifts moved the board over the past year. First, Microsoft folded its Azure OpenAI story into the broader Azure AI Foundry platform and widened the catalog well beyond OpenAI — Anthropic’s Claude models landed there, joining Llama, Mistral, Grok, and DeepSeek. Second, OpenAI’s open-weight GPT OSS models appeared on Amazon Bedrock, and in June 2026 those models — along with NVIDIA Nemotron — received FedRAMP High and DoD IL-4/5 approval in AWS GovCloud. Third, Google kept extending Vertex’s Model Garden while its FedRAMP High authorization stayed in progress rather than GA.

    The net effect: feature parity is now assumed, and the platforms compete on distribution rights and compliance paper. That is a very different evaluation than the one most enterprises ran in 2024.

    Model exclusivity is the real differentiator

    Start with the exclusives, because they are the constraint you cannot negotiate around. OpenAI’s frontier models remain available only through Azure and OpenAI’s own API — no amount of AWS or Google spend commitment gets you GPT-class frontier models on Bedrock or Vertex. The open-weight GPT OSS variants on Bedrock are useful, but they are not the flagship line. Likewise, Gemini is served through Vertex and Google’s own developer API, and Amazon’s Nova models live on Bedrock. If a specific frontier model is a hard requirement, that requirement picks your cloud for you.

    Anthropic is the exception that proves the rule. Claude runs on Bedrock, on Vertex, and on Foundry — with the caveat that feature velocity differs. Anthropic’s newest API capabilities ship on its first-party API and its Anthropic-operated AWS offering first; the partner-operated clouds carry a subset that catches up on their own schedules. The same lag pattern applies to Gemini features between Google’s own API and third-party surfaces. Budget engineering time for those gaps.

    On raw breadth, Vertex leads. Model Garden curates 50+ models spanning Gemini, Claude, Google’s open-weight Gemma line, Llama, and Mistral — the widest first-party frontier selection any single console offers. Foundry’s raw catalog is large and growing, but its center of gravity is still OpenAI. Bedrock’s lineup is the most Claude-centric of the three. Note that this same exclusivity fight plays out one layer up in the productivity suite — see our Copilot vs. Gemini Enterprise comparison for that side of the ledger.

    Why it matters: compliance decides for regulated shops

    For federal agencies, defense contractors, and a growing share of financial services and healthcare buyers, the model comparison is academic until the compliance question clears. Here the gap is stark. Amazon Bedrock holds FedRAMP High, and the June 2026 GovCloud expansion extended High and IL-4/5 coverage to additional models on the platform. Azure OpenAI likewise holds FedRAMP High. Vertex AI’s FedRAMP High authorization, as of mid-2026, is still not generally available.

    Read that plainly: a CIO subject to FedRAMP High has a two-horse race today, no matter how well Gemini benchmarks. Google will presumably close this gap — but “presumably” is not a control you can show an auditor. The takeaway for the meeting: compliance authorizations are point-in-time facts, and you buy what is authorized now, not what is promised.

    Billing traps to model before you commit

    Two Gemini-specific traps deserve a line item in any Vertex evaluation. First, Gemini’s reasoning models bill their internal thinking tokens as output tokens. A hard question that returns a 500-token answer can generate several thousand tokens of hidden reasoning, and you pay output rates for all of it. Teams that estimate cost from visible response length will underestimate — sometimes by multiples.

    Second, Gemini Pro uses tiered long-context pricing: once a single prompt crosses roughly 200K tokens, input and output bill at the higher long-context rates — approximately double the standard tier as of mid-2026 (list pricing varies; check current rate cards). Big-context RAG patterns that casually stuff documents into the prompt will cross that line without anyone noticing until the invoice arrives.

    Bedrock and Foundry have their own cost sharp edges — provisioned throughput commitments on Bedrock and capacity quotas on Foundry’s frontier models chief among them — but they are capacity-shaped, not token-shaped, and easier to see coming. Whatever platform you pick, benchmark the fully loaded per-request cost in your proof of concept. If that number climbs past your comfort line at production volume, run the math against running your own inference — our self-hosted LLM vs. API cost analysis covers the break-even points.

    Where each platform falls short

    • Bedrock: no OpenAI frontier models and no Gemini, so shops standardized on either are out. Newer model-vendor API features often arrive on first-party APIs before Bedrock exposes them, and the developer experience still trails the other two consoles.
    • Vertex AI: the FedRAMP High gap disqualifies it for regulated US workloads today, Gemini’s billing model punishes sloppy prompt engineering, and Google remains the minority enterprise cloud — you may be adding a third cloud relationship just for AI.
    • Azure AI Foundry: capacity quotas on the flagship OpenAI models remain a recurring operational complaint, the broader catalog is newer and less proven than Vertex’s, and the anchor vendor sells the same frontier models direct — Azure exclusivity is a distribution deal, not a technology moat.

    What to do about it

    Four rules of thumb. One: if you carry FedRAMP High or DoD IL requirements, shortlist Bedrock and Azure OpenAI, and revisit Vertex only when its authorization goes GA. Two: absent a compliance gate or a hard model exclusive, default to the cloud where your data and identity already live — cross-cloud AI adds latency, egress, and a second security review for marginal benefit. Three: build your baseline workloads on a model you can move — Claude’s presence on all three clouds makes it the natural portability hedge — and keep prompts and evals in your own repo, not the vendor’s studio. Four: if data sovereignty is the real driver rather than FedRAMP, weigh the platforms against a private deployment; our private AI on VMware vs. OpenShift AI brief covers that path.

    The one-liner for the steering committee: the platforms converged, so buy on model rights and compliance paper — and model Gemini’s token billing before you sign anything.

    Frequently asked questions

    Is Vertex AI FedRAMP High authorized?

    Not as of mid-2026. Vertex AI’s FedRAMP High authorization is in progress but not generally available, which rules it out for US federal workloads and many regulated-industry buyers today. Bedrock and Azure OpenAI both hold FedRAMP High.

    Can I run OpenAI models on AWS Bedrock?

    Only the open-weight GPT OSS variants, which are available on Bedrock and received FedRAMP High and IL-4/5 approval in GovCloud in June 2026. OpenAI’s proprietary frontier models remain exclusive to Azure and OpenAI’s own API.

    Which platform offers the most models?

    Vertex AI, on curated frontier breadth — Model Garden spans 50+ models including Gemini, Claude, Gemma, Llama, and Mistral. Azure AI Foundry’s raw catalog is larger by count, but its frontier depth centers on OpenAI plus recent additions like Claude and Grok.

    Why do Gemini responses cost more than the visible token count suggests?

    Gemini’s reasoning models generate internal thinking tokens that are billed as output even though you never see them. A short answer to a hard question can bill as thousands of output tokens, so estimate cost from measured usage, not response length.

    Is Claude available on all three clouds?

    Yes. Anthropic’s Claude models are served on AWS Bedrock, Google Vertex AI, and Azure AI Foundry, making Claude the most portable frontier model family — though feature availability differs by platform and typically lands on Anthropic’s first-party API first.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Retiring VPN for ZTNA in 2026: A Phased Migration Playbook

    Retiring VPN for ZTNA in 2026: A Phased Migration Playbook

    Roughly 65% of enterprises say they plan to replace their traditional VPN with zero trust network access, and the ZTNA market is projected to grow from about $1.34 billion in 2025 to $4.18 billion by 2030. Those numbers tell you where the industry is going. They do not tell you how to get there, because every vendor publishes only its own half of the story — the onboarding guide for its product, never the 12 to 24 months when VPN and ZTNA run side by side and your help desk fields tickets for both.

    This is the missing middle: a phased playbook for a VPN to ZTNA migration that survives contact with legacy applications, plus honest mechanics for the three platforms most enterprises shortlist — Zscaler Private Access, Palo Alto Networks Prisma Access, and Microsoft Entra Private Access.

    The honest starting point: VPN won’t hit zero

    Every ZTNA pitch deck ends with the VPN concentrator in a dumpster. Real migrations don’t. Legacy applications that lack modern authentication — thick clients speaking proprietary protocols, server-initiated connections, unauthenticated printer and scanner traffic, that ERP module nobody has touched since 2014 — will resist brokered, identity-centric access. Plan for a residual VPN footprint of 5–15% of your application estate for years, or budget the application-remediation work to eliminate it.

    The takeaway for the steering committee: the goal is not “zero VPN.” The goal is shrinking the flat-network blast radius from everything to a short, named list of exceptions with an owner and a review date. Frame it that way on day one and the project survives its first awkward legacy app. Frame it as VPN elimination and you will be explaining a “failed” migration to the board in 18 months.

    Phase 1 — Inventory apps and their auth

    You cannot broker access to applications you cannot name. Start with 4–8 weeks of discovery: pull VPN logs, NetFlow, and firewall data to build the real application list — not the CMDB fiction. Most enterprises find far more internal apps than they expected. Then classify each app on two axes:

    • Auth readiness — does it speak SAML/OIDC, can it sit behind a proxy with header-based auth, or is it hard-coded to NTLM, Kerberos-only, or nothing at all?
    • Traffic pattern — client-initiated web or TCP (easy), long-lived sessions and VoIP (harder), server-initiated or peer-to-peer like remote support tools and softphones (hardest; check your vendor’s connector support explicitly).

    The major platforms now help here — Zscaler’s AI-powered app discovery builds segmentation policy suggestions from observed traffic, and Microsoft’s Quick Access mode lets you onboard broad IP ranges first and tighten to per-app policy later. Use the tooling, but keep a human owner per application. Discovery output is your migration backlog; treat it like one.

    Phase 2 — Run ZTNA and VPN in parallel

    Parallel running is where migrations quietly die, so set rules before you start. Deploy the ZTNA client alongside the VPN client to a pilot ring of 5–10% of users — include your loudest power users deliberately, because they find breakage fast. Route only pilot apps through the broker; everything else stays on VPN. Critically, make the routing decision in the client, not the user: if users have to choose which tunnel to use, they will choose wrong and blame the new thing.

    Two thresholds worth stealing. First, hold each app in parallel for two full business cycles — usually two months, long enough to catch month-end batch behavior — before cutting VPN routes to it. Second, cap the parallel period at 24 months total. Past that, you are paying double licensing, double client overhead, and double help-desk training indefinitely — the migration has become a lifestyle.

    Phase 3 — Migrate app-by-app, not user-by-user

    The instinct is to migrate by department. Resist it. Migrating user-by-user means every user needs every app working on day one — one broken legacy app rolls back the whole cohort. Migrating app-by-app means an application is either served by the broker for everyone or it isn’t, and your exception list shrinks visibly week over week. It also forces the conversation that matters: for each app, either it gets a connector and a policy, it gets remediated to modern auth, or it goes on the residual-VPN exception list with a named owner.

    Sequence by risk and ease: start with internal web apps behind SSO (fast wins, visible progress), then TCP thick clients, then the awkward tail. And write policy for non-human identities as you go — service accounts, RPA bots, and increasingly AI agents need scoped access paths too, a problem we unpack in our AI agent identity governance guide. Bolting that on after the human migration means doing discovery twice.

    Phase 4 — Decommission at roughly 80% coverage

    Do not wait for 100%. Once roughly 80% of internal applications are served through the broker, start pulling VPN infrastructure: shrink concentrator capacity, cut licensing tiers at renewal, and move the residual VPN to a segmented enclave that reaches only the exception-list apps — not the flat network. That last step is the one enterprises skip, and it is the whole point. A legacy VPN that lands users in a five-app enclave is a contained risk; one that lands them on the old flat network means you ran a two-year project without changing your blast radius.

    The VP-ready takeaway: 80% brokered coverage is the decommission trigger, and the residual VPN must terminate in an enclave, not the core.

    Vendor mechanics: Zscaler, Palo Alto, Microsoft

    Zscaler Private AccessPrisma Access (Palo Alto)Entra Private Access (Microsoft)
    ModelCloud broker; Client Connector plus outbound App ConnectorsSASE fabric with firewall-grade inspection on the access pathIdentity-first; Conditional Access extended to private apps
    Strongest atMature broker at scale; app discovery and segmentationInline threat inspection of private-app traffic; PAN-OS shopsEntra ID-standardized orgs; licensing already in the suite
    Watch forPremium pricing; platform lock-in as scope growsModule and licensing sprawl; highest typical TCOYounger product; deepest value assumes a Microsoft stack
    Best fitLarge enterprises with aggressive segmentation goalsSecurity-first shops wanting one inspection stack everywhereM365-heavy mid-to-large orgs optimizing spend

    Zscaler Private Access

    ZPA is the most battle-tested pure broker: outbound-only App Connectors mean nothing is exposed inbound, and its AI-assisted user-to-app segmentation genuinely shortens Phase 1. Where it is weaker: costs climb as you extend coverage from remote workers to office users and third parties, and you are committing to Zscaler’s cloud as the control plane for everything. Best fit for large enterprises that want segmentation depth and can absorb premium pricing.

    Prisma Access (Palo Alto Networks)

    Prisma Access’s differentiator is that private-app traffic gets the same threat prevention stack — intrusion prevention, malware analysis, DNS security — as internet-bound traffic, and its ZTNA Connector handles app-side connectivity without inbound holes. The trade-off is complexity: multiple modules, fragmented licensing, and typically the highest total cost of the three, though Panorama-managed shops recover much of that in operational familiarity. We compared the two SASE heavyweights directly in our Zscaler vs Palo Alto SASE brief.

    Microsoft Entra Private Access

    Microsoft’s entry, part of the Global Secure Access family, extends the Conditional Access policies you already run for SaaS to private applications, and it has moved fast — B2B guest access and Intelligent Local Access shipped in late 2025, with browser-based BYOD access in preview as of early 2026. It is the value play if you already pay for Entra Suite or the newer M365 bundles that include it. It is also the youngest product here: inspection depth and non-Windows client maturity trail the specialists, as of mid-2026. Microsoft-centric organizations should shortlist it first; heterogeneous shops should pilot it hardest.

    Frequently asked questions

    Can ZTNA completely replace a VPN?

    For most enterprises, not entirely — and not soon. Client-initiated web and TCP applications migrate cleanly; server-initiated traffic, legacy auth, and some VoIP or thick-client workloads do not. Plan for a 5–15% residual VPN footprint confined to a segmented enclave, and treat every app on that list as remediation debt with a named owner.

    How long does a VPN to ZTNA migration take?

    Budget 12–24 months for a mid-to-large enterprise: one quarter of discovery, one quarter of pilot, then app-by-app waves. Under 12 months usually means the inventory was skipped; past 24 months of parallel running means you are paying for two access stacks indefinitely and should force the exception-list decision.

    Which applications don’t work well with ZTNA?

    The usual offenders: server-initiated connections (remote support, some management tools), peer-to-peer traffic like softphones, unauthenticated device traffic such as printing and scanning, and anything hard-wired to legacy network auth. Vendor support varies — verify your specific protocols against each platform’s connector documentation before you commit to a decommission date.

    Is ZTNA cheaper than a VPN?

    Per-user list pricing is usually higher than VPN licensing, and you pay double during the parallel phase. The savings arrive later: retired concentrators, a smaller breach blast radius, fewer lateral-movement incidents, and — for Microsoft-licensed shops — ZTNA effectively bundled into suites you already buy. Model it over three years, not one; list pricing varies enough that you should negotiate all three.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.