David Davis

  • Kubernetes Backup in 2026: Veeam Kasten vs. Rubrik

    Kubernetes Backup in 2026: Veeam Kasten vs. Rubrik

    Somewhere in the last three years, Kubernetes backup stopped being a Velero-and-a-prayer side project and became a real budget line. The reason is stateful workloads: once databases, message queues, and now virtual machines run on the cluster, “we can redeploy from Git” stops being a recovery strategy — etcd state and persistent volumes have to be protected like any tier-1 system. Two vendors dominate the enterprise shortlist for that job: Veeam Kasten, the purpose-built Kubernetes-native platform, and Rubrik, the cyber-resilience vendor that treats backup as a security control first.

    I have sat through enough data-protection renewals to know these two rarely lose to anyone else — they lose to each other. This brief gives you the verdict, the side-by-side, the honest downsides of both, and the OpenShift wrinkle that is quietly deciding more of these deals than either vendor admits.

    The verdict

    If Kubernetes is your primary platform and you run it across multiple distributions and clouds, buy Veeam Kasten. If ransomware recovery is the board-level driver and you want Kubernetes folded into a single Zero Trust data-security platform alongside VMs, SaaS, and databases, buy Rubrik. That’s the short version, and it holds in most evaluations I’ve watched.

    The nuance: these products come from opposite directions. Kasten was built inside Kubernetes — it discovers applications by namespace and label, understands operators and CRDs, and treats the cluster as the unit of truth. Rubrik extends an enterprise-wide security architecture down into Kubernetes — immutable, append-only storage, air-gapped vault options, and clean-room recovery workflows that assume the attacker is already inside. Neither approach is wrong. They just answer different questions.

    Side-by-side comparison

    Veeam KastenRubrik
    Design centerKubernetes-native; runs in the cluster, app-aware by namespace, label, and operatorSecurity-first; Kubernetes as one workload inside Rubrik Security Cloud
    Auto-discoveryPolicy-driven discovery of new apps and namespaces; new workloads inherit protectionSLA-domain model auto-applies policy to discovered Kubernetes resources
    ImmutabilitySupported via object-lock targets (S3 Object Lock and equivalents)Structural — append-only immutable filesystem, Zero Trust access, air-gap and vault options
    Ransomware recoveryRestore points, encryption, KMS integration; recovery tooling improvingClean-room recovery, anomaly detection, threat hunting across backups
    Distribution coverageBroadest as of mid-2026: OpenShift, EKS, AKS, GKE, Rancher, Tanzu, vanilla upstreamMajor distributions and managed clouds; narrower validated matrix
    VM storyKubeVirt / OpenShift Virtualization with incremental, changed-block-aware backupOpenShift Virtualization protected alongside VMware and Nutanix in one platform
    LicensingPer-node subscription, licensed separately from Veeam Data PlatformCapacity/subscription within the broader Rubrik platform; list pricing varies
    Best fitPlatform teams, multi-cluster and multi-cloud estates, OpenShift Virtualization adoptersCISO-driven buyers consolidating data security across VMs, SaaS, and Kubernetes

    Where Veeam Kasten wins

    Kasten K10 — rebranded Veeam Kasten for Kubernetes after the acquisition, and still licensed separately from Veeam’s core suite — remains the deepest Kubernetes-native data protection product on the market. It keeps pace with upstream aggressively (Kubernetes 1.34 support landed quickly), and its coverage matrix is the widest in the category: OpenShift, EKS, AKS, GKE, Rancher, and vanilla clusters all get first-class treatment. If your estate spans more than one distribution or managed-Kubernetes service, that breadth alone can settle the evaluation.

    The operational model is the other differentiator. Kasten’s policies auto-discover new applications and namespaces, so a workload deployed on Tuesday is protected Tuesday — no ticket, no backup admin in the loop. In 2026 that policy-driven auto-discovery is table stakes for any serious buyer; Kasten simply does it with more Kubernetes fluency than anyone else, including application-consistent hooks for databases and awareness of operators and custom resources. Its KubeVirt support has also matured into a genuine strength: incremental, changed-block-aware VM backup developed with Red Hat’s storage-agnostic APIs, which matters enormously for the wave of VMware refugees now running VMs on OpenShift Virtualization.

    Takeaway for the meeting: Kasten is the choice when Kubernetes is the platform, not a workload.

    Where Rubrik wins

    Rubrik’s pitch starts from a different premise: assume breach. Everything protected by Rubrik Security Cloud lands in an append-only, immutable filesystem with Zero Trust access controls — immutability is structural, not a checkbox you configure on a bucket. On top of that sits the security tooling that has made Rubrik a board-room name: anomaly detection on backup data, threat hunting across restore points, and clean-room recovery workflows built around Rubrik Cloud Vault for rebuilding from a known-clean copy after a ransomware event. The company was named a Leader in the 2026 Gartner Magic Quadrant for Backup and Data Protection Platforms, and its momentum is real.

    For Kubernetes specifically, Rubrik delivers application-centric protection of objects and persistent volumes under the same SLA-domain policy model it uses for everything else. That consistency is the point. If your organization already runs Rubrik for VMware, databases, and M365, adding Kubernetes means one console, one policy language, one recovery playbook — and one throat to choke during an incident. Rubrik has also extended coverage to OpenShift Virtualization, so the VMs-on-Kubernetes crowd is no longer forced elsewhere.

    Takeaway: Rubrik is the choice when the CISO owns the budget and recovery-from-attack is the scenario being funded.

    The honest downsides

    Veeam Kasten

    • It is a separate product with separate licensing — owning Veeam Data Platform does not give you Kasten, which surprises procurement teams and pushes real-world cost above the sticker impression.
    • Immutability depends on the target you point it at. Object-lock done right is solid, but the burden of getting it right sits with your team, not the product’s architecture.
    • The security analytics layer — anomaly detection, threat hunting — is thinner than Rubrik’s. Kasten protects data well; it does less to tell you the data is compromised.

    Rubrik

    • Kubernetes depth trails Kasten’s. Distribution coverage is narrower, and cluster-native nuances — operators, complex CRD-based apps, exotic CSI drivers — get less first-class handling.
    • The structural security comes with structural opinion. You buy into Rubrik’s platform, its storage model, and its pricing; flexibility is the trade for the Zero Trust posture.
    • Buying Rubrik only for Kubernetes rarely pencils out. The economics work when it consolidates several workload types; as a point solution it is expensive.

    The OpenShift factor

    Red Hat is the third vendor in this comparison whether it wants to be or not. OpenShift Virtualization has become the leading landing zone for enterprises exiting VMware after the Broadcom repricing, and every one of those migrations drags a VM backup requirement onto Kubernetes. Red Hat ships OADP (its Velero-based operator) for basic protection, and it is genuinely fine for small estates — but at enterprise scale, with hundreds of VMs and application-consistency requirements, most teams graduate to a commercial platform within the first year.

    That graduation is where this comparison gets decided. Kasten’s Red Hat engineering collaboration on changed-block tracking for KubeVirt gives it the more efficient per-VM backup path today; Rubrik counters by protecting OpenShift Virtualization inside the same platform that still covers your remaining VMware estate — a cleaner story mid-migration. Rule of thumb: if the VMware exit will finish inside 18 months, weight Kasten; if you will run dual hypervisor estates longer than that, Rubrik’s single pane earns its premium.

    What to do about it

    • Score your driver first. Platform coverage and Kubernetes fluency → Kasten. Ransomware resilience and audit posture → Rubrik. Deals that skip this step run six months long.
    • Demand auto-discovery in the PoC. Deploy a new namespace mid-trial and verify it is protected without human action. Both vendors claim this; make them show it on your clusters.
    • Test the restore, not the backup. Time a full application restore — PVs, config, secrets — into a different cluster. Cross-cluster recovery is where the two products’ philosophies actually diverge on the stopwatch.
    • Price the immutable copy explicitly. Object-lock storage for Kasten, vault capacity for Rubrik — the clean-copy line item is where quotes converge more than list pricing suggests.

    If you are earlier in the process and still building requirements, start with our Kubernetes backup buyer’s guide before you invite either vendor in — the RFP questions matter more than the logo you pick.

    Frequently asked questions

    Is Veeam Kasten included with Veeam Backup & Replication?

    No. Veeam Kasten for Kubernetes is licensed separately from the Veeam Data Platform, typically per worker node. Budget for it as its own line item, and confirm node counts across all clusters — including dev and staging if you intend to protect them.

    Is Velero good enough for enterprise Kubernetes backup?

    For small clusters and stateless-heavy estates, often yes. Once you have stateful tier-1 applications, compliance-driven immutability requirements, or VMs on KubeVirt, the operational cost of scripting around Velero usually exceeds a commercial license within a year.

    Does Rubrik support Red Hat OpenShift Virtualization?

    Yes. Rubrik added OpenShift Virtualization protection to Rubrik Security Cloud, with the same immutable, policy-driven model it applies to VMware and Nutanix — a deliberate play for the post-VMware migration wave.

    What are the main Kasten K10 alternatives in 2026?

    Rubrik is the strongest commercial alternative, followed by Portworx (data-management-led), Trilio, and CloudCasa. Red Hat’s OADP covers basic OpenShift needs. Shortlist by driver — security posture, storage integration, or price — rather than feature checklists.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Buy or Rent GPUs? The 2026 Break-Even Math

    Buy or Rent GPUs? The 2026 Break-Even Math

    A single NVIDIA B200 lists somewhere between $30,000 and $40,000 as of mid-2026 — and the same GPU rents for roughly $4 to $7 per hour depending on where you shop. Divide one number by the other and you get the question every AI budget meeting now turns on: at what point does renting silicon stop making sense? I have sat through enough of these CFO-meets-CTO sessions to know the answer is not “it depends.” It is a utilization number, and you can calculate it this quarter.

    This brief lays out the 2026 break-even math, where AWS and Google Cloud actually price Blackwell capacity, what the buy path really costs once you count lead times, and a decision rule you can defend in front of finance. One warning before we start: most of what ranks in search for this question is written by GPU clouds selling one side of the answer. Treat their calculators accordingly.

    What changed

    Blackwell went from paper launch to rentable reality. AWS now sells EC2 P6-B200 instances — eight B200s per node with 1,440 GB of GPU memory — in multiple US regions, plus P6e-GB200 UltraServers for rack-scale Grace Blackwell. Google Cloud’s A4 VMs put the same HGX B200 board on-demand at roughly $4.30 per GPU-hour at list. Meanwhile H100-class capacity, the workhorse of 2024–25, has become a buyer’s market: specialty clouds advertise it under $2/hour while hyperscaler list pricing still sits in the $4–8 range.

    The other change is on the buy side. Blackwell supply improved through 2026, but enterprise orders still quote 6–12 month lead times for full systems once you include networking, power provisioning, and integration. That lag is a cost, and most buy-vs-rent models quietly leave it out.

    The break-even math

    The verdict first: at sustained utilization of roughly 60–70%, owning beats renting on a three-year horizon. Below that, rent. The table shows the shape of the market as of mid-2026 — list pricing varies, so run your own numbers with your negotiated rates.

    Rent — specialty GPU cloudRent — hyperscalerBuy — on-prem or colo
    H100-class, per GPU-hour~$1.50–3.00~$4–8 (list)~$1.30–2.00 effective at 70%+ utilization
    B200-class, per GPU-hour~$4–7 on-demand~$4.30 (Google Cloud A4 list) and up~$2.20–2.80 effective at full utilization
    Upfront capitalNoneNone (commits optional)$30–40K per B200 GPU + facility costs
    Time to capacityHours to daysHours (quota permitting)6–12 months for full systems
    Best fitBursty training, experimentsData-gravity and compliance-adjacent workSteady-state inference, 24/7 pipelines

    Here is the arithmetic behind the crossover. A B200 at roughly $35K amortized over three years is about $1.33 per hour of raw depreciation. Add power, cooling, networking, colo space, and operations staff and the fully loaded figure lands near $2.20–2.80 per hour — if the card is busy every hour. At 60% utilization that effective cost climbs to roughly $3.70–4.60, which is exactly where Blackwell rental pricing sits. That is the break-even. Our on-prem GPU cluster cost guide walks through the facility-side line items most spreadsheets miss.

    Why it matters

    Because the default is drifting. In 2024 the safe answer was “rent everything — the hardware cycle is too fast to own.” In 2026 that reflex is costing real money for anyone running production inference around the clock. An inference fleet at 80% utilization on rented hyperscaler capacity can pay for equivalent owned hardware in under 18 months. Boards have noticed, and finance teams are asking why AI compute is 100% opex when the workload profile looks like a steady utility.

    The reverse mistake is just as expensive. Teams that bought H100 clusters in 2024 for “future training needs” and ran them at 25% utilization effectively paid $6–8 per GPU-hour for capacity they could have rented for $3. Utilization, not unit price, decides this argument. The takeaway a VP can repeat: rent your spikes, own your baseline.

    Renting: AWS, Google Cloud, and the neoclouds

    AWS is the depth play. P6-B200 instances slot into the same VPC, IAM, and data-platform surface your teams already run, and Capacity Blocks let you reserve GPU windows for defined training runs instead of paying on-demand rates around the clock. The catch is price: AWS GPU list pricing runs well above specialty clouds on a per-GPU-hour basis, and the gap only closes with Savings Plans or negotiated commits. If your training data already lives in S3 and your security team has blessed the account structure, that premium buys real friction reduction. If not, you are paying hyperscaler rates for undifferentiated silicon.

    Google Cloud is currently the sharpest hyperscaler pencil on Blackwell. A4 VMs at roughly $4.30 per GPU-hour at list undercut comparable AWS on-demand pricing meaningfully, and Dynamic Workload Scheduler plus spot capacity can push effective rates lower for interruptible training. Google also gives you an escape hatch NVIDIA-only shops lack: TPU capacity for workloads that fit it. The weakness is the familiar one — smaller enterprise footprint, and quota negotiations that favor large committed spenders. Google Cloud fits teams that treat compute as a market and shop it; it fits less well where the org is contractually welded to another cloud.

    Below both sit the neoclouds — CoreWeave, Lambda, and a long tail — renting H100-class capacity at $1.50–3.00 per hour. Real savings, real trade-offs: thinner enterprise support, variable data-center quality, and contract terms that reward diligence.

    Buying: the NVIDIA hardware path

    NVIDIA sets the terms on the buy side, and its interests are not subtle: it sells to you, to AWS, to Google, and to every neocloud simultaneously. A B200 at $30–40K list is only the opening line item. HGX baseboards, NVLink switching, InfiniBand or Ethernet fabric, and 10kW+ per-node power budgets typically push a deployed cluster to 1.6–2x the GPU line. NVIDIA’s strength for buyers is the software moat — CUDA, NIM, and enterprise support make owned hardware genuinely productive on day one. The weakness is that you are buying at the top of a fast product cadence: Blackwell Ultra and the Rubin generation are already on NVIDIA’s public roadmap, which compresses the resale value of whatever you rack today.

    Model the 6–12 month lead time as a cost, not a footnote. If you must rent B200 capacity at $5/hour while your purchased cluster clears procurement, facilities, and bring-up, a 64-GPU stopgap can add seven figures to the “buy” column before your hardware serves a single token. For inference-heavy shops, the buy decision also interacts with the build-vs-API question — our self-hosted LLM vs API cost analysis covers when owning the model layer pays.

    Honest downsides on both sides

    • Renting: price volatility (Blackwell on-demand rates have moved double-digit percentages within a year), quota ceilings at exactly the moment everyone wants capacity, egress charges on training data, and the quiet ratchet of commits that turn “flexible opex” into a three-year contract anyway.
    • Buying: capital tied up in a depreciating asset on a roughly 18-month product cadence, lead-time exposure, the staffing reality that a GPU cluster needs datacenter and MLOps skills you may not have, and utilization risk — the whole case collapses if the workload you bought for gets cancelled.

    What to do about it

    Instrument first. Pull 90 days of actual GPU utilization before any procurement conversation. Then apply the rule: workloads sustained above ~70% utilization go on owned or colo hardware; anything below ~50% stays rented; the 50–70% band is your negotiation zone for reserved capacity and committed-use discounts. Rent Blackwell for training bursts and evaluation runs. Buy — or lease through colo — for the inference baseline you can forecast two years out. Re-run the math every two quarters, because rental pricing in this market does not sit still. And when a vendor’s calculator tells you their side wins, check whose logo is on the spreadsheet.

    Frequently asked questions

    Is it cheaper to buy or rent GPUs for AI?

    It depends on utilization, not preference. As of mid-2026, owning wins on a three-year horizon once sustained utilization passes roughly 60–70%. Below that, rental pricing — especially H100-class capacity under $3/hour — is hard to beat.

    How much does it cost to rent a B200 GPU per hour?

    On-demand B200 pricing spans roughly $4–7 per GPU-hour as of mid-2026. Google Cloud’s A4 VMs list near $4.30 per GPU-hour; specialty clouds and spot capacity can run lower, hyperscaler on-demand can run higher.

    How much does an NVIDIA B200 cost to buy?

    List pricing runs roughly $30–40K per GPU, but a deployed cluster — baseboards, fabric, power, cooling, integration — typically lands at 1.6–2x the GPU line item, with 6–12 month lead times for full systems.

    Should I use AWS or Google Cloud for GPU workloads?

    Google Cloud currently prices Blackwell more aggressively at list; AWS offers deeper enterprise integration and reservation mechanics like Capacity Blocks. If your data and security posture already live on one of them, that gravity usually outweighs the list-price gap.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Self-Hosted LLM vs. API in 2026: The Real Cost per Token

    Self-Hosted LLM vs. API in 2026: The Real Cost per Token

    Every AI platform review I have sat in this year hits the same slide: “should we still be paying per token?” The math finally has real answers. As of mid-2026, a 70B-class open model served from an 8x H100 pod lands in the range of $1–$2.50 per million tokens at healthy utilization — FP8 quantization pushes the low end near $1 — while frontier API list prices for comparable quality run anywhere from $0.10 to $15 per million depending on tier. Those two ranges overlap, and that overlap is exactly why so many teams get this decision wrong.

    This brief gives you the break-even bands, the honest per-million numbers on both sides, and the costs that never make it into the spreadsheet. Short version: volume and utilization decide this, not ideology.

    The break-even numbers

    Here is the framework I give every architecture team that asks. Treat the boundaries as bands, not lines — caching behavior, model class, and traffic shape move them — but the bands themselves have held up across every deployment I have reviewed this year.

    Monthly token volumeCheapest option, as of mid-2026Why
    LowUnder ~20M tokensManaged APIIdle GPU hours dominate; even one rented H100 at ~$3/hr is money burned at this volume
    Gray zone~20M–100M tokensIt dependsPrompt caching, batch discounts, and whether a small model suffices decide the winner
    High100M+ tokens, sustainedSelf-hosted or dedicated rented GPUs$1–$2.50/M on an 8x H100 pod undercuts frontier list pricing — if utilization stays high

    The single most important variable is utilization. At roughly 10% utilization, your real cost per token can run near 10x the headline rate — an idle H100 is more expensive per token than a premium frontier API. Self-hosting wins on volume you actually serve, not volume you project in a planning deck.

    Where the self-hosting math comes from

    The self-hosting story in 2026 is still overwhelmingly an NVIDIA story. Cloud H100 pricing has stabilized around $2.85–$3.50 per hour at most providers after falling hard from the 2024 peaks, and the software side — TensorRT-LLM, the NIM inference microservices, and community stacks like vLLM running on NVIDIA silicon — is where the per-token gains actually come from. FP8 quantization on Hopper-class hardware is the difference between roughly $4.50 per million tokens for a 70B model and something closer to $1–$1.70. That is NVIDIA’s real strength here: the optimization ecosystem is mature and the talent pool knows it.

    The weakness is equally clear: you are buying into a supply-constrained, premium-priced platform, and the depreciation math is brutal. Blackwell-class hardware makes a pod you bought eighteen months ago look slow, and resale values reflect it. Whether you should own the metal at all is a separate decision — our GPU buy-vs-rent break-even analysis covers that math, and for teams going on-prem, the full cluster cost guide itemizes the power, cooling, and networking lines that double the sticker price.

    Takeaway for the meeting: self-hosting a 70B-class model at $1–$2.50 per million tokens is real, but only above sustained high utilization on hardware someone still has to buy, rack, or reserve.

    The API counterargument: prices mostly fell

    The strongest argument against self-hosting is that the API vendors keep cutting the ground out from under your business case. Since mid-2025, effective per-token prices across the major providers have fallen on the order of 40–60% for comparable capability — through cheaper tiers, aggressive caching discounts, and batch pricing.

    Indicative list price (per M in / out)Where it fits
    Google Gemini 2.5 Flash-Lite~$0.10 / $0.40High-volume classification, extraction, routing
    Google Gemini 2.5 Flash~$0.30 / $2.50General workhorse tier
    OpenAI nano tier~$0.10–$0.20 inputCheapest branded-model volume play
    Anthropic Sonnet-class~$3 / $15Mid-frontier reasoning; flat-rate long context

    List pricing varies and changes fast, so treat these as shapes, not quotes. Each vendor’s shape is different. Google’s Gemini line is the volume-pricing aggressor — Flash-Lite at roughly $0.10 in / $0.40 out is hard for any self-hosted small model to beat once you count engineering time. But note the direction is not uniformly down: Google raised Gemini 2.5 Flash list pricing during 2025 when it consolidated tiers, a useful reminder that a business case built on someone else’s price list carries repricing risk in both directions.

    OpenAI competes hardest at the cheap end — its nano-tier models undercut nearly everything for simple high-volume tasks, plus roughly 90% off cached input reads and a 25% batch discount. Anthropic is rarely the cheapest per token; its case is quality per dollar at the mid-frontier tier and flat-rate pricing on very long context windows, which matters if your workload is document-heavy. Both Anthropic and OpenAI now discount cached prompt reads by about 90%, which quietly demolishes naive break-even math: if 70% of your tokens are a cached system prompt, your effective API rate is a fraction of list.

    Takeaway: never compare self-hosting against API list prices. Compare it against your blended effective rate after caching and batching — that number is often 3–5x lower.

    The middle path: managed open models

    The binary framing — own GPUs or pay OpenAI — misses where a lot of enterprises actually land. AWS is the clearest example of the middle path: Bedrock gives you per-token access to Anthropic and open-weight models with enterprise controls and no infrastructure, SageMaker lets you serve open models on capacity you rent but do not rack, and EC2 P5-class instances with Capacity Blocks cover the full self-managed route. AWS also fields its own Inferentia silicon, which can undercut NVIDIA per-token for supported models — at the cost of a smaller software ecosystem and porting work.

    The honest critique of the middle path is margin stacking: a managed open model on Bedrock costs more per token than the same model on hardware you operate well. You are paying AWS to hold the operational risk. For many teams in the 20M–100M gray zone, that is a rational trade. How the three hyperscaler platforms compare on exactly this is covered in our Bedrock vs. Vertex AI vs. Azure Foundry breakdown.

    The costs nobody puts in the spreadsheet

    A “free” open-weight model is the most expensive line item in some AI budgets I have reviewed. The recurring omissions:

    • Engineering headcount. Serving infrastructure, quantization, eval pipelines, and on-call for a production inference stack realistically consumes 2–4 senior engineers — $500K+ per year fully loaded before you serve a single token.
    • Model churn. Open-weight leaders change every few months. Every migration re-runs evals, red-teaming, and capacity planning.
    • Failover capacity. APIs bundle redundancy into the price. Self-hosters buy it twice.
    • Peak-to-average ratio. You provision for peak; you pay for average. Bursty traffic quietly halves your effective utilization — and doubles your real per-token cost.

    What to do at your volume

    • Under 20M tokens/month: stay on APIs. Spend the energy on prompt caching and model right-sizing instead — that is where your 40% savings actually is.
    • 20M–100M/month: run a 90-day bake-off. Price your real blended API rate after caching, then quote dedicated rented capacity for an open model that passes your evals. Do not buy hardware yet.
    • 100M+/month sustained: self-hosting almost always wins on unit cost — if you can keep utilization above roughly 60% and staff the platform team. Rent before you buy, and hold the depreciation risk consciously.

    The sentence for the meeting: below 20 million tokens a month the API is cheaper than your own idle GPUs, above 100 million self-hosting wins, and everyone in between should be negotiating — with both sides.

    Frequently asked questions

    Is it cheaper to self-host an LLM than use an API?

    Only above sustained volume — roughly 100M+ tokens per month with 60%+ GPU utilization. Below about 20M tokens per month, managed APIs are almost always cheaper once idle hardware and engineering time are counted.

    How much does it cost to run a 70B model per million tokens?

    As of mid-2026, roughly $1–$2.50 per million tokens on 8x H100-class capacity at healthy utilization, with FP8 quantization near the low end. Poor utilization can multiply that by 5–10x.

    Do LLM API prices keep going down?

    Mostly. Effective prices fell roughly 40–60% since mid-2025 through cheaper tiers, caching, and batch discounts — but not uniformly. Google raised Gemini 2.5 Flash list pricing in 2025, so build repricing risk into any multi-year case.

    What are the hidden costs of self-hosting LLMs?

    Engineering headcount ($500K+/year for a credible platform team), redundant failover capacity, model migration and re-evaluation every few months, and the gap between peak provisioning and average utilization.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Zscaler vs. Palo Alto SASE in 2026: How to Pick

    Zscaler vs. Palo Alto SASE in 2026: How to Pick

    Two vendors sit on almost every enterprise SASE shortlist in 2026, and they are no longer variations of the same product. Zscaler is the pure cloud-proxy zero-trust bet: all traffic rides its exchange, and secure access is the entire company. Palo Alto Networks is the platform-consolidation bet: SASE as one arm of a firewall, SOC, and browser estate run from one vendor. Pick by feature checklist and you miss the point — you are choosing an architecture and an operating model for the next five years, and the two roads diverge more now than they did in 2023.

    This brief gives you the verdict up front, the side-by-side, the honest downsides of both, the pricing shape buyers actually see, and a decision framework you can defend in the budget meeting.

    The verdict

    Buy Zscaler when secure access is the product you are actually buying. It is the cleaner architecture for retiring VPN, it works with whatever firewalls and EDR you already own, and its proxy cloud is the most battle-tested in the category. Buy Prisma SASE when you already run a substantial Palo Alto firewall estate, when Cortex is your SOC direction, or when unmanaged devices are a first-class problem — its natively integrated secure browser is something Zscaler does not have as of mid-2026. Neither choice is wrong. One of them is wrong for your environment.

    What changed

    Gartner split the market view, and the split is informative. Zscaler remains a Leader in the Magic Quadrant for Security Service Edge — its fourth consecutive year in that position — while landing as a Visionary in the newer single-vendor SASE Platforms Magic Quadrant, where the scoring rewards owning the SD-WAN and branch stack outright. Palo Alto Networks is the only vendor named a Leader in all three years of that SASE Platforms MQ. Read together: Zscaler leads the security-service layer; Palo Alto leads the converged single-vendor platform story.

    Both vendors also moved. Palo Alto shipped Prisma SASE 4.0 and pushed Prisma Browser — its enterprise secure browser — past six million licensed seats by late 2025, turning BYOD and contractor access into a genuine differentiator. Zscaler spent its money on the SOC side of the house, most visibly the Red Canary acquisition, a move we unpack in our analysis of Zscaler’s agentic SOC ambitions. Both are converging on the same thesis: secure access is the delivery vehicle for a much larger platform sale.

    Side-by-side comparison

    ZscalerPalo Alto Prisma SASE
    ArchitecturePurpose-built cloud proxy (Zero Trust Exchange)Cloud-delivered firewall stack plus SD-WAN and browser
    Analyst positionLeader, SSE MQ (4 straight years); Visionary, SASE Platforms MQLeader, SASE Platforms MQ (3 straight years); Leader, SSE MQ
    Unmanaged devicesCloud browser isolation; no native enterprise browserPrisma Browser, natively integrated, 6M+ licensed seats
    Branch / SD-WANZero Trust SD-WAN — newer, thinner offeringPrisma SD-WAN — five-time SD-WAN MQ presence
    SOC integrationBuilding via acquisition (Red Canary)Cortex XSIAM, mature and deeply tied in
    Effective pricingRoughly $8–25/user/month by module mixComparable band; heavily shaped by platform bundling
    Best fitVendor-neutral shops retiring VPN at scaleExisting Palo Alto firewall and Cortex estates

    Where Zscaler wins

    Zscaler’s advantage is focus. The Zero Trust Exchange was built as a multi-tenant proxy cloud from day one, not adapted from firewall software, and it shows in operational maturity: ZIA for internet and SaaS traffic and ZPA for private application access are the reference implementations most competitors get measured against. If your driving project is getting off VPN concentrators — and for most enterprises in 2026 it is — ZPA is the shortest path, a migration we lay out step by step in our VPN-to-ZTNA playbook.

    The second advantage is neutrality. Zscaler does not care whose firewalls sit in your data center or whose EDR runs on your endpoints. For enterprises with heterogeneous estates — a Fortinet branch layer here, CrowdStrike there, a Cisco campus — that neutrality keeps the SASE decision from forcing three other decisions. Takeaway for the meeting: Zscaler is the strongest pick when secure access must stand on its own merits rather than lean on an existing vendor relationship.

    Where Palo Alto wins

    Palo Alto wins on convergence. If your firewall estate is already NGFW and Panorama, Prisma Access extends the same policy model to remote users — one policy language, one management plane, one renewal. The pull gets stronger if Cortex XSIAM is your SOC direction, because access telemetry lands natively in the same detection pipeline. This is the same consolidation gravity reshaping cloud security, which we covered in our look at the Google–Wiz CNAPP consolidation — buyers are trading best-of-breed for fewer throats to choke, and Palo Alto is built to collect that trade.

    The sharpest single differentiator is Prisma Browser. Contractors, M&A onboarding, BYOD clinicians — populations where you cannot install an agent — get enterprise-grade controls inside a managed browser instead of through clunky VDI. Zscaler’s answer is cloud browser isolation, which solves a narrower problem at a different price-performance point. If a third of your workforce is unmanaged, this line item alone can decide the deal.

    The honest downsides

    Zscaler first. Module sprawl is real: what starts as ZIA quietly becomes ZIA plus ZPA plus ZDX plus data protection, and the per-user math climbs toward the top of the range. Once all traffic transits the exchange, switching costs are high and renewal leverage sits with the vendor — negotiate accordingly. The branch story is the thinnest part of the portfolio; Zero Trust SD-WAN is young, and complex branch estates will still want a dedicated SD-WAN vendor alongside.

    Palo Alto’s downsides are the mirror image. Licensing is genuinely hard to model — platformization deals bundle SASE with firewall refreshes and Cortex commitments in ways that obscure unit economics until renewal. Prisma Access has historically trailed Zscaler on cloud-native operational polish, and admins coming from the firewall world face a steeper console learning curve than Zscaler’s single-purpose UI. And the consolidation that makes the deal attractive is also the trap: you are deepening dependence on one vendor across network, access, and SOC simultaneously.

    What to do about it

    Three rules of thumb hold up in real procurements. First: if more than roughly 60 percent of your firewall estate is already Palo Alto and Cortex is your stated SOC direction, Prisma SASE is the default and Zscaler must beat it on proof, not promise. Second: if you are vendor-neutral and VPN retirement is the funded project, Zscaler is the default for the opposite reason. Third: if unmanaged users exceed about 20 percent of your access population, weight Prisma Browser heavily — replicating it on the Zscaler side means a separate product decision.

    • Run a 500-user proof of concept on production traffic for 30 days — measure p95 latency to your top ten SaaS apps from your three worst geographies, not the vendor’s demo regions.
    • Price the three-year total, not year one — as of mid-2026, list pricing varies widely and discounting is aggressive for multi-module, multi-year commitments on both sides.
    • Refuse shelfware: buy only the modules you will deploy inside 12 months, and lock pricing for the ones you might add later.

    The repeatable line for the steering committee: Zscaler is the best secure-access product; Palo Alto is the best secure-access platform. Decide which one you are actually buying.

    Frequently asked questions

    Is Zscaler better than Palo Alto?

    Neither is categorically better. Zscaler leads on proxy-based zero trust and vendor-neutral deployment; Palo Alto leads on single-vendor convergence, SD-WAN depth, and unmanaged-device coverage via Prisma Browser. Your existing estate usually decides it.

    What is the difference between Zscaler and Prisma Access?

    Zscaler is a purpose-built cloud proxy exchange sold as a standalone security service. Prisma Access is Palo Alto’s cloud-delivered firewall stack, designed to share policy and management with on-prem NGFWs and the broader Prisma SASE and Cortex portfolio.

    How much does Zscaler cost per user?

    Effective pricing runs roughly $8–25 per user per month as of mid-2026 depending on module mix — internet access alone sits at the low end, while full bundles with private access, digital experience monitoring, and data protection reach the top. List pricing varies; multi-year deals discount heavily.

    Is Palo Alto a Leader in SASE?

    Yes — Palo Alto Networks is the only vendor named a Leader in all three years of Gartner’s Magic Quadrant for SASE Platforms, and it is also a Leader in the Security Service Edge MQ, where Zscaler has led four consecutive years.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Nutanix External Storage: AHV Without the HCI Rip-and-Replace

    Nutanix External Storage: AHV Without the HCI Rip-and-Replace

    For a decade, the most reliable objection in any Nutanix evaluation had nothing to do with the hypervisor. It was the array. “We have three years left on our SAN — moving to HCI means writing it off.” I have watched that single line kill more AHV proposals than any technical gap ever did. As of mid-2026, the objection is mostly dead: Nutanix now runs compute-only nodes against Pure Storage FlashArray over NVMe/TCP, Dell PowerFlex is supported, NetApp ONTAP and Dell PowerStore are in early access, and January’s FlashStack with Nutanix launch turned the Pure option into a pre-validated converged stack from Cisco, Pure, and Nutanix.

    This brief covers what actually shipped, what it does to the VMware-exit math, and who should still buy integrated Nutanix nodes anyway.

    What changed

    Nutanix built its brand on collapsing three-tier infrastructure. Every node ran AOS, compute and storage scaled together, and your existing array had no role except the recycling dock. That architecture was the pitch — and the problem, because it forced a storage buy into every deal.

    Through 2025 and into 2026, Nutanix deliberately walked that back. Nutanix Cloud Infrastructure now supports compute-only clusters that consume external arrays: Pure Storage FlashArray (//X, //XL, and //C models) over NVMe/TCP is generally available, Dell PowerFlex is supported through its native SDC client on PowerEdge servers, and NetApp ONTAP — announced at .NEXT 2026 as an NFS-based integration — is in early access with general availability targeted for the second half of 2026. Dell PowerStore sits in early access as well.

    Then on January 13, 2026, Cisco, Pure Storage, and Nutanix announced general availability of FlashStack with Nutanix: UCS compute, FlashArray storage, and NCI with AHV as a jointly validated converged architecture, with lifecycle management stitched across Intersight, Prism, and Pure1. That is not a compatibility footnote. It is three vendors co-selling a packaged alternative to vSphere on converged infrastructure.

    Why it matters for the VMware-exit math

    Every VMware-exit business case I have seen since the Broadcom acquisition has the same three ugly lines: new hardware, data migration, and project risk. External storage support attacks all three.

    • Hardware: an array with two or three years of depreciation and support left keeps doing its job. You buy compute-only nodes, not a full HCI refresh.
    • Migration: the data does not move. You present the same array to AHV compute-only nodes, migrate VMs off ESXi, and retire hosts. Zero-data-migration is the phrase that changes steering-committee body language.
    • Risk: a hypervisor swap with storage in place is a materially smaller project than a hypervisor swap plus a storage platform swap. Fewer moving parts, shorter rollback path.

    The verdict is short: the single largest capex objection in the Nutanix column of your Nutanix vs. VMware cost model just left the spreadsheet. If you priced an AHV migration in 2024 and rejected it on hardware grounds, that number is stale. Reprice it.

    The external storage matrix, as of mid-2026

    ProtocolStatus (mid-2026)Notes
    Pure Storage FlashArrayNVMe/TCPGA//X, //XL, //C qualified; the FlashStack with Nutanix reference design
    Dell PowerFlexPowerFlex SDC clientGARuns on PowerEdge; software-defined, not NVMe/TCP
    NetApp ONTAPNFSEarly accessAnnounced at .NEXT 2026; GA targeted second half of 2026
    Dell PowerStoreNVMe/TCPEarly accessTimeline not yet firm — plan accordingly

    Read the status column literally. Early access means you can pilot, not that you can build a Q4 production migration on it. And every row carries qualifications — supported server platforms, protocol constraints, and feature dependencies — so treat the hardware compatibility list as a gate, not a suggestion.

    The vendor-by-vendor read

    Pure Storage: first mover, deepest integration

    Pure got there first and it shows: NVMe/TCP end to end, three FlashArray families qualified, and the only external option wrapped in a jointly validated converged design. If your arrays are already Pure, this is the shortest credible path off ESXi. The trade-offs are the usual ones — FlashArray carries premium list pricing, and the FlashStack design assumes you are willing to standardize on the Cisco-Pure pairing. Pure’s Evergreen consumption model, though, dovetails neatly with the whole premise here: keep the array, upgrade in place, stop forklifting.

    Dell: the biggest installed base, the most mixed incentives

    Dell brings two tracks. PowerFlex is generally available today, but it is a software-defined system using Dell’s own SDC client rather than NVMe/TCP, and it expects PowerEdge underneath — fine if you are a Dell shop, irrelevant if you are not. PowerStore, the array most mid-size enterprises actually own, is still early access. Be clear-eyed in account conversations: Dell also sells competing stacks and has its own hypervisor-adjacent interests, so the enthusiasm of your Dell team for a Nutanix attach may vary. Push for roadmap dates in writing.

    NetApp: the sleeper, if the date holds

    The ONTAP integration announced at .NEXT 2026 matters because of who owns ONTAP: a huge share of exactly the enterprises now staring at Broadcom renewal quotes. NFS-based consumption keeps it operationally simple. The caution is timing — early access now, GA targeted for the second half of 2026. If your renewal lands before the GA date plus a sensible pilot window, ONTAP support is a 2027 lever, not a 2026 one.

    Cisco: the packaging play

    Cisco’s contribution is UCS compute and the validated-design muscle that makes FlashStack with Nutanix a one-motion purchase — reference architecture, joint support posture, Intersight lifecycle management. For shops already standardized on UCS, that removes most of the integration homework. Worth remembering: after HyperFlex’s retirement, Cisco needed a hyperconverged-adjacent answer, and this is it. Cisco sells the compute either way — which is precisely why the partnership is stable.

    Nutanix: a reversal, and a smart one

    Credit where due — Nutanix spent years arguing three-tier was legacy, and now it sells into three-tier. That is pragmatism, and it widens the funnel enormously. The honest caveat: on external arrays, data services live on the array. Compression, dedupe, snapshots, and replication become the array vendor’s job, and some AOS-native capabilities do not apply. The “one platform, one support call” story gets more nuanced the moment a second storage vendor is in the room.

    Who should still buy integrated Nutanix nodes

    External storage support does not kill the HCI pitch; it narrows it to where it was always strongest. Buy integrated nodes when the array is within roughly twelve months of a support renewal or refresh anyway, when you are deploying edge or ROBO sites where a SAN never made sense, or when you want one operational model and one throat to choke — AOS data services, native DR, and a single upgrade train. That consolidation value is real; it was just never worth writing off a healthy array to get.

    My rule of thumb: two or more years of array support runway and adequate headroom, go compute-only against the array. Renewal inside twelve months, price both paths — the integrated node quote will often be closer than you expect once array support renewal costs land in the model.

    What to do about it

    • Inventory array support runway and utilization before any vendor conversation. The array’s remaining life is now a negotiating asset in every VMware alternatives evaluation, not a constraint.
    • Ask Nutanix to quote compute-only nodes against your existing array and full HCI nodes side by side. Make them show both numbers.
    • Check feature dependencies early: metro clustering, synchronous replication, and backup vendor integration behave differently on external storage than on AOS. Get specifics per array, in writing.
    • Pilot on a non-production cluster before committing a migration wave — especially anything currently in early access.
    • If you are also weighing keeping vSphere but dropping vSAN, run the same array-reuse logic across the vSAN alternatives field; the pattern is identical.

    The takeaway for the meeting: Nutanix no longer requires a storage decision to make a hypervisor decision. Price your exit accordingly.

    Frequently asked questions

    Does Nutanix AHV support external SAN storage?

    Yes. As of mid-2026, Nutanix compute-only clusters can consume Pure Storage FlashArray over NVMe/TCP and Dell PowerFlex in general availability, with NetApp ONTAP (NFS) and Dell PowerStore in early access. Support comes with qualified hardware lists and protocol requirements, so verify your exact configuration.

    What is FlashStack with Nutanix?

    A converged infrastructure architecture from Cisco, Pure Storage, and Nutanix, generally available since January 2026. It pairs Cisco UCS servers and Pure FlashArray with Nutanix Cloud Infrastructure and AHV, validated jointly and managed through Intersight, Prism, and Pure1.

    Do I lose Nutanix data services with external storage?

    Some. On external arrays, capacity efficiency and replication are handled by the array rather than AOS, and certain AOS-native features do not apply to compute-only deployments. Map your DR and backup requirements against the specific array integration before committing.

    Is moving to Nutanix cheaper than renewing VMware?

    It depends on your renewal quote, host counts, and hardware timing — but external storage support removes the forced array write-off that used to sink the comparison. Model compute-only nodes against your existing array before assuming the incumbent renewal wins.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Phishing Click Rate Benchmarks 2026: KnowBe4 Data by Industry

    Phishing Click Rate Benchmarks 2026: KnowBe4 Data by Industry

    One in three. That is how many untrained employees click a simulated phish, according to KnowBe4’s 2026 Phishing by Industry Benchmarking Report — a 33.2% global baseline Phish-prone Percentage drawn from 42 million simulations across 14.8 million users at roughly 64,000 organizations. If your board asks whether the security awareness line item is working, this is the dataset you benchmark against. This brief gives you the 2026 numbers by industry and organization size, the thresholds that separate a healthy program from a stalled one, and the honest caveats about what simulation click rates do and do not tell you.

    The 2026 benchmarks

    The headline numbers first. KnowBe4 measures Phish-prone Percentage (PPP) — the share of users who click a link, open an attachment, or otherwise fail a simulated phishing test. The 2026 report covers 19 industries, four organization sizes and seven global regions.

    Phish-prone Percentage (2026)
    Global baseline, untrained (all industries)33.2%
    Healthcare & Pharmaceuticals, untrained42.7%
    Insurance, untrained38.1%
    Retail & Wholesale, untrained36.0%
    Large enterprise (10,000+ employees), untrained39.5%
    After 12 months of sustained training (all industries)~4.2%

    Two things stand out. First, size hurts: large enterprises of 10,000-plus employees start at 39.5%, meaningfully worse than the global average, because scale brings more turnover, more contractors and more inboxes nobody owns. Second, the trained number is the story. Organizations that run continuous simulated phishing and training for a full year land around 4.2% — a reduction of roughly 87% from baseline. That delta is the strongest ROI evidence anywhere in the security awareness category.

    How to read your own click rate

    Benchmarks only matter if they change a decision. Here is how I read a phishing click rate benchmark by industry when a team brings me their numbers:

    • Above 30% on your first baseline test: normal. Do not panic and do not punish anyone — you are simply average, and the fix is a program, not a memo.
    • Still above 15% after six months of training: your program is broken. Either the cadence is too slow (quarterly is not a cadence), the templates are too easy to game, or training is a once-a-year compliance video.
    • Between 5% and 10% after a year: respectable but unfinished. The residual clickers are usually concentrated — new hires, specific departments, shared mailboxes. Segment the data before spending more.
    • Under 5% sustained: mature. Shift budget from volume of simulations to depth — harder spear-phish templates, callback phishing, QR codes and MFA-fatigue scenarios.

    The takeaway for the meeting: a good phishing click rate in 2026 is under 5%, and anything above 15% a year into a funded program is a program failure, not a people failure.

    Why healthcare, insurance and retail keep losing

    Healthcare & Pharmaceuticals tops the vulnerability table again at 42.7%, and in large healthcare organizations the untrained rate climbs even higher. The reasons are structural, not cultural. Clinical staff work under time pressure on shared workstations, email is a life-or-death coordination channel that people process fast, and turnover keeps the untrained population permanently refreshed. Insurance (38.1%) and Retail & Wholesale (36%) share a related profile: large distributed workforces, heavy seasonal hiring, and daily legitimate email that looks exactly like phishing bait — invoices, claims, shipping notices, password resets.

    If you run security in one of these three industries, the benchmark is your budget argument. You are starting from a measurably worse position than the cross-industry average, and breach economics compound the problem — as our 2026 data breach cost benchmarks show, healthcare remains the most expensive industry in which to get breached. A higher click rate feeding a higher cost-per-breach is exactly the kind of multiplication a CFO understands.

    What the training curve actually looks like

    The 87% reduction is real, but it is not fast. KnowBe4’s longitudinal data shows the biggest gains arrive between month three and month twelve — not in the first 90 days. That has two practical implications. First, a pilot program judged at 90 days will look mediocre and may get cut exactly when it is about to start working. Set expectations with leadership for a 12-month evaluation window. Second, one-off annual campaigns do not bend the curve at all. The organizations hitting 4.2% run continuous simulations — typically at least monthly — with immediate, short remedial training at the moment of failure.

    Worth noting: the year-over-year trend is favorable. The equivalent 2025 report showed an 86% reduction from training; 2026 shows roughly 87%. The methodology is consistent enough across years that the direction is credible, and it means the intervention keeps working even as attackers adopt AI-generated lures.

    KnowBe4’s role — and where the data has limits

    KnowBe4 owns this benchmark for a simple reason: nobody else has the sample size. With tens of thousands of customer organizations feeding the dataset, its annual report is the closest thing the industry has to a census of simulated phishing behavior. The company has also moved decisively beyond the awareness-training label — its HRM+ platform bundles simulated phishing and training with real-time coaching (SecurityCoach), cloud email security, and a growing suite of AI agents under the AIDA banner, eight of them as of early 2026. That repositioning mirrors the broader market shift we covered in the move from security awareness to human risk management: the product category is no longer annual training, it is continuous measurement and intervention on human risk.

    Where KnowBe4 is strong: breadth of template and content library, mature automation for continuous campaigns, and benchmark data that makes board reporting easy — you can put your PPP next to your industry’s line and be done. Where it is weaker: per-seat pricing adds up at large-enterprise scale, the sheer content volume can overwhelm smaller teams without a program owner, and a simulation-centric metric invites gaming — run easy templates and your numbers look great while your real risk does not move.

    And treat the dataset itself with analyst discipline. It is drawn from KnowBe4’s own customer base, which skews toward organizations that already bought awareness tooling. PPP measures clicks on simulations, not real-world compromise — a useful proxy, not a ground truth. And the 33.2% baseline is measured before training by definition, so the dramatic headline delta is partly a property of how the cohort is constructed. None of that invalidates the numbers. It does mean you should benchmark against the trend and your industry line, not treat 4.2% as a guarantee.

    What to do about it

    • Baseline before budget season. Run an unannounced simulation across the full employee population and put your number next to the 33.2% global line and your industry’s line. That one slide anchors the whole funding conversation.
    • Commit to twelve months, not a pilot. The curve bends between months three and twelve. Fund a year of at least monthly simulations with in-the-moment remedial training, and report quarterly against the 87% reduction trajectory.
    • Segment the stragglers. Once you are under 10%, stop treating the workforce as one population. New hires, finance, executive assistants and shared mailboxes deserve targeted, harder scenarios.
    • Audit template difficulty annually. If your click rate dropped but your reported-phish rate did not rise, you are probably measuring easier templates, not better behavior.

    Frequently asked questions

    What is a good phishing click rate?

    Under 5% on realistic simulations, sustained over multiple campaigns, is a mature result in 2026. Between 5% and 10% is acceptable for a program in its first year. Above 15% after a year of funded training signals a broken program cadence.

    What is the average phishing simulation click rate in 2026?

    KnowBe4’s 2026 benchmarking data puts the untrained global average at 33.2%, falling to roughly 4.2% after twelve months of continuous simulation and training. Healthcare (42.7%), insurance (38.1%) and retail (36%) start above the average.

    What does Phish-prone Percentage mean?

    Phish-prone Percentage (PPP) is KnowBe4’s metric for the share of users who fail a simulated phishing test — clicking a link, opening an attachment, or submitting data. It measures simulation behavior, which is a proxy for real-world susceptibility rather than a direct measure of compromise.

    How much does security awareness training reduce phishing clicks?

    Across KnowBe4’s 2026 dataset, organizations running sustained programs cut susceptibility by about 87% within a year — from a 33.2% baseline to around 4.2%. Most of the improvement lands between month three and month twelve, which is why one-off campaigns underperform.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Kubernetes Backup in 2026: Kasten, Portworx, or Velero?

    Kubernetes Backup in 2026: Kasten, Portworx, or Velero?

    Most Kubernetes backup evaluations I see in 2026 start the same way: a team installs Velero because it’s free, runs its first test restore, hits a CSI VolumeSnapshot error, and only then starts pricing commercial tools. That sequence wastes a quarter. The market has consolidated to three real options — Veeam Kasten, Portworx PX-Backup, and Velero — and the right answer depends on your storage, your staffing, and increasingly on whether you’re running virtual machines on Kubernetes at all.

    This guide gives you the verdict up front, an honest look at where each tool falls down, the KubeVirt VM wrinkle that VMware refugees keep tripping over, and the one prerequisite check that derails more first deployments than anything else.

    The short version

    Kasten is the default. Buy Portworx if you’re a Pure Storage shop already running Portworx Enterprise. Run Velero only if you have platform engineers with time to own it. As of mid-2026, practitioner surveys put Kasten’s mindshare at roughly a third of the Kubernetes backup market — about triple PX-Backup’s — and that gap shows up in hiring pools, community answers, and integration coverage.

    Veeam KastenPortworx PX-BackupVelero
    Cost modelCommercial, per-node licensing; free tier for small clustersCommercial, strongest value bundled with Portworx EnterpriseFree and open source; you pay in engineering hours
    KubeVirt VM backupFirst-class: VM discovery, label-based policies, CBT work with Red HatStrong when VMs sit on Portworx storage; live migration supportPossible via OADP on OpenShift, but assembly required
    Storage dependencyStorage-agnostic (any CSI driver with snapshots)Best on Portworx storage; works elsewhere with caveatsStorage-agnostic via CSI or plugins
    Ops burdenLow — policy engine, UI, RBAC built inLow-to-moderateHigh — CLI-driven, you build the guardrails
    FitsMost enterprises, especially Veeam and OpenShift shopsPure Storage / Portworx storage customersSmall estates, non-prod, strong platform teams

    Veeam Kasten: the default for a reason

    Kasten K10 became Veeam Kasten for Kubernetes after the acquisition, and it has spent the years since compounding its lead. The policy engine is genuinely application-aware — it captures namespaces, secrets, CRDs, and volume data as a unit, and its blueprint system handles databases that need quiescing before a snapshot. On the VM side it discovers KubeVirt virtual machines automatically and treats them as workloads, and label-based VM policies mean a new VM tagged into a protection group is covered the moment it’s created. Kasten and Red Hat have also co-developed a storage-agnostic changed-block-tracking approach for KubeVirt, which addresses the ugliest problem in VM-on-Kubernetes backup: full-volume reads on every backup pass.

    The honest downsides: per-node licensing gets expensive at scale, and pricing conversations at renewal reflect Kasten’s market position. It also remains a separate product with a separate console from Veeam Data Platform — convergence is coming, with native OpenShift Virtualization support signaled for the v13-era roadmap, but today you run two panes of glass. If you’re weighing Kasten against Rubrik’s Kubernetes protection specifically, we’ve done that comparison in depth in our Kasten vs Rubrik brief.

    Portworx PX-Backup: strongest inside a Pure shop

    Pure Storage owns Portworx, and that ownership defines the buying logic. If your clusters already run Portworx Enterprise as the storage layer, PX-Backup is the path of least resistance: container-granular, application-consistent backups that understand the storage underneath them, plus live migration for KubeVirt VMs on RWX block volumes. For OpenShift Virtualization on Portworx storage, the combined story — storage, VM mobility, and backup from one vendor — is coherent in a way few competitors match.

    Outside that context, the case weakens. PX-Backup standalone against Kasten on third-party storage is a harder sell, and its roughly 10% mindshare means fewer engineers arrive knowing it. The rule of thumb: Portworx storage in production makes PX-Backup a finalist; no Portworx storage means it rarely wins the bake-off.

    Velero: free, with an operational tax

    Velero is everywhere, and for good reason — it’s free, it’s GitOps-friendly, and it underpins Red Hat’s own OADP operator. For a handful of clusters, or for non-production estates, it’s a defensible choice. The tax comes due at scale: no real multi-tenant RBAC, a plugin compatibility matrix you own, restores you must test yourself because nobody else will, and upgrades that occasionally break snapshot workflows. There is no vendor to call at 2 a.m.

    My threshold: above roughly ten production clusters, or any compliance regime that requires demonstrable restore testing, the engineering hours spent babysitting Velero exceed the cost of a commercial license. Below that line, Velero plus disciplined restore drills is fine. A VP can defend either position — but only if the line was drawn deliberately.

    The VMware refugee problem: VMs on Kubernetes

    Here’s what makes 2026 different. Teams leaving VMware — and there are many; see our guide to VMware alternatives in 2026 — are landing on OpenShift Virtualization in real numbers, which means the backup tool now has to protect KubeVirt VMs, not just containers. A VM is not a stateless pod. It needs crash-consistent or better snapshots, incremental capture so you’re not re-reading a 2 TB disk nightly, and restore workflows an ops team can actually run.

    Kasten currently has the most mature answer, as covered above. But it’s not alone. Rubrik has added OpenShift Virtualization protection to its platform — a meaningful move for shops that already standardized on Rubrik for ransomware resilience and want one policy plane across VMs old and new (our Veeam vs Rubrik comparison covers how the two philosophies differ). Trilio deserves a look too: it’s a smaller vendor whose entire business is cloud-native data protection, with application-centric backups spanning namespaces, labels, and VMs, and an OpenShift and OpenStack heritage that plays well in telco and Red Hat-heavy environments. Trilio’s risk profile is the inverse of its focus — you’re betting on a specialist. And Veeam’s signaled v13-era native OpenShift Virtualization support means VBR shops may eventually get VM-on-Kubernetes backup inside the console they already run.

    The CSI snapshot prerequisite that derails deployments

    Every tool in this guide — commercial or free — depends on CSI VolumeSnapshots for persistent volume capture, and this is where most first deployments die. The failure is silent: backups appear to succeed, then the first restore returns empty volumes. Before you install anything, verify four things:

    • Your CSI driver actually supports snapshots — not all do, and hostPath or in-tree drivers never will.
    • The external-snapshotter CRDs and snapshot controller are installed. Several managed distributions still ship without them.
    • A VolumeSnapshotClass exists and is annotated as default for the driver.
    • A manual VolumeSnapshot on a test PVC completes and is restorable — prove the plumbing before you buy the appliance.

    Plan the backup target at the same time. All three tools export to S3-compatible object storage, and on-prem teams should size that tier deliberately — our on-prem S3 object storage guide walks through the options.

    What to do about it

    • Running Portworx Enterprise storage? Shortlist PX-Backup first, Kasten second.
    • OpenShift Virtualization in the migration plan? Weight KubeVirt maturity heavily — Kasten leads today, with Rubrik and Trilio as credible alternatives depending on your incumbent stack.
    • Fewer than ten production clusters and a capable platform team? Velero is defensible — budget the engineering hours honestly and schedule quarterly restore drills.
    • Everyone else: start with Kasten, negotiate the node count, and run the CSI snapshot checklist before the proof of concept, not during it.

    Frequently asked questions

    Is Velero good enough for production Kubernetes backup?

    For small estates with a strong platform team, yes — provided you test restores on a schedule. Past roughly ten production clusters, or under audit regimes requiring documented recovery evidence, the operational cost usually exceeds a commercial license.

    What happened to Kasten K10?

    It still exists — Veeam renamed it Veeam Kasten for Kubernetes. Same product line, now positioned inside the broader Veeam portfolio, with a free tier remaining for small clusters.

    How do I back up VMs running on OpenShift Virtualization?

    Use a tool that treats KubeVirt VMs as first-class workloads. Veeam Kasten discovers VMs automatically and supports label-based policies and changed-block tracking; Rubrik and Trilio both protect OpenShift Virtualization as well. Plain Velero via OADP works but requires significant assembly.

    Why do my Kubernetes backups succeed but restores come back empty?

    Almost always a CSI snapshot problem: the snapshot controller or CRDs are missing, or no default VolumeSnapshotClass exists, so the tool captured manifests but no volume data. Verify a manual VolumeSnapshot works before trusting any backup job.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Can Your Storage Array Detect Ransomware? NetApp vs Pure

    Can Your Storage Array Detect Ransomware? NetApp vs Pure

    In March 2026, NetApp — the vendor with the loudest claim to storage-native ransomware detection — signed alliances with Commvault and Elastio, two companies whose entire business is inspecting data after it leaves the array. That is the storage industry quietly answering the question buyers keep asking me: if my array can detect ransomware, do I still need detection anywhere else? The answer, from the vendor best positioned to say no, is yes.

    This brief lays out what NetApp’s ARP/AI actually catches, what Pure Storage’s SafeMode does instead, why the two approaches are not interchangeable, and how the Commvault and Elastio deals settle the detection-versus-immutability argument. If you are budgeting a ransomware line item for FY27, this is the framing to take into the meeting.

    What changed

    Three things moved this from a feature-checklist debate to a strategy question. First, NetApp’s Autonomous Ransomware Protection with AI (ARP/AI) became the first storage-native detection engine to earn SE Labs’ AAA rating, posting 99% detection accuracy with zero false positives on legitimate workloads in that test. Second, NetApp wrapped it in a Ransomware Recovery Guarantee — a warranty-style commitment, with configuration strings attached, that snapshot recovery will work. Third, and most telling: in March 2026 NetApp announced partnerships with Commvault and with Elastio, embedding Elastio’s deep file inspection of snapshots into its Ransomware Resilience Service and wiring ONTAP’s detection signals into Commvault’s recovery workflows.

    The takeaway a VP can repeat: the vendor with the best array-level detection just paid for two more layers of it. That is not an admission of failure — it is an admission of scope.

    Why it matters

    Ransomware economics turn on two clocks: time to detect and time to recover clean. Endpoint and network tooling owns the first clock until the attacker reaches your data. Once encryption starts hitting primary storage, the array is the last observer with a real-time view — and the first place a fast response can shrink the blast radius from terabytes to gigabytes. A detection that fires minutes into an encryption run, paired with an automatic snapshot, is the difference between restoring a volume and rebuilding a data center.

    But detection at the array is inherently a file-workload story. ARP/AI watches NAS activity — entropy shifts, extension churn, abnormal write patterns on NFS and SMB shares. Attackers who exfiltrate quietly, encrypt inside application-level containers, or simply go after your backups first never trip that wire. That is why the immutability camp, led by Pure, has argued for years that guaranteed-clean copies matter more than clever alarms. Both camps are half right, which is exactly what the March deals confirm.

    NetApp’s bet: detect at the array

    ARP/AI runs on-box in ONTAP, using a pre-trained model that needs no per-workload learning period — a genuine improvement over the original ARP, which required weeks of baselining before it could be trusted in enforcement mode. When it suspects an attack it takes an automatic snapshot and raises an alert, so your recovery point lands minutes into the incident rather than at last night’s backup. The SE Labs result — AAA, 99% detection, zero false positives in testing on ONTAP 9.15.1 — is the strongest independent validation any storage vendor has for on-box detection, and NetApp has been justifiably loud about it.

    Where it is weaker: ARP/AI is a file-workload feature. SAN and block-heavy estates get far less benefit, and a lab result on curated samples is not a guarantee against a patient adversary who throttles encryption below detection thresholds. The Recovery Guarantee is real money but reads like an insurance policy — specific configurations, specific processes, or no payout. NetApp fits shops that are already ONTAP-centric with large unstructured estates, and it is the obvious pick when file-share ransomware is your top tabletop scenario. For the broader platform question, see our full Pure Storage vs. NetApp comparison.

    Pure’s bet: snapshots nobody can delete

    Pure Storage — mid-rebrand to Everpure as of mid-2026, though the product names have not moved — takes the opposite position: assume detection fails, and make the copies indestructible. SafeMode retention locks snapshots so that nobody, including a fully compromised storage admin account, can delete them before the timer expires — up to 30 days on FlashArray, up to 400 days on FlashBlade. Unlocking early requires contacting Pure support with pre-designated named contacts and a multi-step verification process. It ships in the box at no extra license cost, which procurement will notice.

    The honest critique runs the other way: SafeMode is containment, not detection. It tells you nothing while the attack is happening. Pure1’s fleet analytics can flag anomalies — a sudden drop in data reduction ratio is a classic encryption tell — but that is telemetry-scale signal measured in hours, not an on-box engine measured in minutes, and Pure has no SE Labs-style third-party detection rating to point at. The 30-day FlashArray ceiling also matters: dwell times beyond a month are common, and a locked snapshot of already-encrypted data is a very safe copy of garbage. Pure fits block-heavy and mixed estates that want a guaranteed recovery floor with near-zero operational overhead — and teams honest enough to admit they will not tune an alerting pipeline.

    Detection vs. immutability, side by side

    The verdict up front: NetApp wins on detection, Pure wins on simplicity of containment, and neither closes the case alone.

    NetApp ARP/AIPure SafeMode
    Core mechanismOn-box AI detection of encryption behavior on file workloadsRetention-locked snapshots deletion-proofed against admin compromise
    Third-party validationSE Labs AAA, 99% detection, 0 false positives (ONTAP 9.15.1 test)None for detection; immutability is architectural, not tested behavior
    Blast-radius effectShrinks it — auto-snapshot minutes into an attackCaps it — guarantees a floor to recover from
    Coverage gapBlock/SAN workloads, slow-and-low encryptionNo real-time signal; 30-day FlashArray lock vs. longer dwell times
    Operational liftLow-moderate — alerts need an owner and a runbookNear zero — set the policy, size the snapshot capacity
    Best fitONTAP shops, large NAS/unstructured estatesBlock-heavy or mixed estates wanting a guaranteed recovery floor

    One cost note both sides underplay: aggressive snapshot policies consume real capacity, and SafeMode’s lock means a mis-sized policy cannot be walked back quickly. Budget 15–20% headroom before you turn either program on. List pricing varies; neither feature carries a separate license, but the capacity to feed them is not free.

    Does array detection replace backup scanning?

    No — and the March 2026 announcements are the proof. NetApp’s Elastio deal embeds Elastio’s Provable Recovery Control into the NetApp Ransomware Resilience Service, adding deep file inspection of snapshots — actually opening and validating the data, not just watching write patterns — starting with Amazon FSx for NetApp ONTAP. Elastio’s pitch has always been that a snapshot is not a recovery point until something has proven it clean, and NetApp just endorsed that pitch by productizing it. Elastio remains a young company betting on a single capability, which is a vendor-viability question worth asking in diligence, but the capability itself is the missing piece between “we have snapshots” and “we can restore with confidence.”

    The Commvault alliance answers the other half: recovery orchestration. Commvault brings backup-layer anomaly detection, threat scanning, and — critically — the workflow to rebuild at scale, now consuming ONTAP’s detection signals directly. Commvault’s strength is breadth across heterogeneous estates; its weakness is that breadth comes with a heavier operational footprint than either array feature. The pattern to internalize: array detection shrinks the blast radius, immutable copies cap it, and backup-layer inspection proves your way out. Three layers, three different failure modes. Our immutable backup storage comparison covers the second layer across vendors in detail.

    What to do about it

    • If you run ONTAP with meaningful NAS estates, turn ARP/AI on now — it is the rare security feature with independent test results and near-zero workload cost. Assign an alert owner before you enable it, not after.
    • If you run Pure, enable SafeMode this quarter, but size retention against your realistic dwell-time assumption — 14 days is a floor, not a default, and FlashBlade’s 400-day ceiling is where long-retention use cases belong.
    • Whichever array you own, do not cut backup-layer scanning from the budget. Array detection is a tripwire; it is not verification. Rule of thumb: if you cannot prove a recovery point is clean, you do not have one.
    • If your board is asking about worst-case recovery, evaluate an isolated recovery environment as the fourth layer — our cyber recovery vault comparison breaks down who does that well.

    Frequently asked questions

    Can a storage array detect ransomware?

    Yes, for file workloads. NetApp’s ARP/AI detects encryption behavior on NFS and SMB shares in real time and takes automatic snapshots when it fires. Block workloads and slow, throttled encryption remain hard for any array-level engine to see, which is why array detection is a layer, not a strategy.

    Is NetApp ARP/AI better than Pure Storage SafeMode?

    They solve different problems. ARP/AI is real-time detection with strong third-party validation; SafeMode is deletion-proof containment with almost no operational overhead. A NAS-heavy ONTAP shop gets more from ARP/AI; a block-heavy estate gets more from SafeMode. Mature programs want both capabilities, from whichever vendors they run.

    Do I still need backup ransomware scanning if my array has detection?

    Yes. Array detection watches write behavior; backup-layer tools like Commvault and Elastio inspect the data itself and prove recovery points are clean. NetApp’s own 2026 partnerships with both companies signal that even the best array-level detection is one layer in a defense-in-depth design.

    How accurate is storage-level ransomware detection?

    The best independent number as of mid-2026 is SE Labs’ test of NetApp ARP/AI: 99% detection accuracy with zero false positives, earning a AAA rating. Treat lab numbers as an upper bound — real estates include workloads and attacker behaviors no test set covers.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • On-Prem S3 Object Storage After MinIO: A Buyer’s Guide

    On-Prem S3 Object Storage After MinIO: A Buyer’s Guide

    The minio/minio GitHub repository was archived in April 2026, months after the project’s maintainers declared maintenance mode and pointed commercial users toward the proprietary AIStor product. If you run MinIO in production — and tens of thousands of organizations do — the software underneath your backup repositories, data lakes, and CI artifact stores now receives no new features, no compatibility updates, and no guaranteed security patches. Search volume for “MinIO alternatives” tells you how many teams just discovered this the hard way.

    This is a forcing function, not a crisis. Object storage does not stop serving GETs because a repo got archived. But every quarter you wait narrows your options and raises the migration bill. This guide maps the replacement field by use case: commercially supported platforms for production backup targets and Object Lock immutability, and community projects that are fine for labs but not for anything an auditor will ask about.

    What changed

    The timeline matters because it shows this was a strategy, not an accident. In late 2025 MinIO’s leadership announced the open-source project was entering maintenance mode. Features had already been stripped from the community console earlier in the year. The GitHub repository was archived in early 2026, and by April 2026 the archival was final. MinIO the company is very much alive — it now sells AIStor, a supported commercial platform aimed at large-scale AI and data-lake estates. What ended is the free, community-maintained MinIO that most enterprises actually deployed.

    The practical consequence: any MinIO binary you run today is frozen. No CVE response is guaranteed. No compatibility work will track changes in the S3 API, Veeam integration requirements, or Kubernetes releases. The takeaway for the meeting: MinIO didn’t die — the free ride did, and running archived storage software in an auditor-facing role has a shelf life measured in quarters, not years.

    Why on-prem S3 matters now

    Three workloads are driving on-prem S3-compatible object storage in 2026, and none of them tolerate an unsupported platform.

    • Immutable backup repositories. Veeam, Commvault, and Rubrik all treat S3 Object Lock as the standard mechanism for ransomware-resistant backups. An immutable repository that stops receiving security patches is a contradiction in terms — see our immutable backup storage comparison for how the major platforms implement Object Lock.
    • AI data lakes. GPU clusters need local, high-throughput reads. Feeding training and RAG pipelines from cloud object storage means paying egress on every epoch, which is why so many AI estates are repatriating hot datasets.
    • Egress avoidance and sovereignty. If restores, analytics reads, and AI pipelines repeatedly pull the same data, on-prem capacity at a predictable cost per terabyte beats metered egress. Run the math against our backup storage cost-per-TB benchmarks before assuming the cloud is cheaper.

    Rule of thumb: if you re-read more than roughly a third of your stored dataset per month, on-prem object storage usually wins on total cost. Below that, cloud tiers stay competitive.

    The replacement field at a glance

    ModelSweet spotWatch out for
    Scality ARTESCASoftware-defined, commercialImmutable backup targets, 20 TB to multi-PBYou supply and manage the hardware
    Cloudian HyperStoreSoftware or appliance, commercialMixed workloads, multi-tenant, hundreds of PBHeavier platform; real sizing and licensing conversation
    Object First OotbiTurnkey applianceVeeam-only immutable backup, fastest deploymentNot general-purpose S3; Veeam only by design
    MinIO AIStorCommercial successor to MinIOLarge AI/data-lake estates staying on MinIO techPriced for large deployments, not small clusters
    SeaweedFS / Garage / RustFSOpen source, communityLabs, dev/test, hobbyist self-hostingYou are the support organization

    Commercial platforms: Scality, Cloudian, Object First, AIStor

    Scality ARTESCA

    ARTESCA is Scality’s lighter-weight, security-hardened object store, and it is the most direct commercial answer to the post-MinIO backup-target question. It starts around 20 TB on a single node and scales into the petabytes — Scality cites validation with the major backup vendors at 8.5 PB-class configurations. The security posture is the pitch: hardened Linux, MFA, no root shell, and S3 Object Lock in compliance mode for Veeam and Commvault repositories. Where it is weaker: it is software, so you own hardware selection and lifecycle, and for exabyte-scale or heavy multi-tenancy Scality would steer you to its RING product instead. Fit: mid-size to large enterprises that want a supported, hardened backup target without buying a closed appliance.

    Cloudian HyperStore

    Cloudian has sold S3-compatible storage on-prem longer than almost anyone, and its S3 API fidelity is among the best in the field — a real consideration if applications, not just backup software, will hit the endpoint. HyperStore starts at a three-node cluster and grows to hundreds of petabytes by adding nodes, with multi-tenancy, QoS, and billing built in, which is why service providers like it. Cloudian has also leaned hard into the AI story, including GPU-adjacent deployments. The trade-offs: it is a bigger platform with more operational surface than an appliance, and pricing is a negotiated, capacity-based conversation. Fit: large enterprises and providers consolidating backup, archive, and data-lake workloads onto one namespace.

    Object First Ootbi

    Ootbi is a turnkey immutable-backup appliance founded by Veeam’s co-founders and built solely for Veeam. It is the shortest path from purchase order to hardened repository — rack it, connect it to Veeam, and immutability is on by default, with the vendor claiming deployment in about 15 minutes and no storage or security expertise required. That focus is also its boundary: Ootbi is not a general-purpose S3 endpoint for applications or AI pipelines, and cluster capacity is appliance-bounded rather than hundreds-of-petabytes scale. Fit: Veeam shops that want immutability now and would rather buy an outcome than operate a platform. If that is you, the ARTESCA-versus-Ootbi decision is the one to study closely.

    MinIO AIStor

    Staying with MinIO commercially is a legitimate option, and for some estates the right one. AIStor is the supported successor, aimed at exascale AI and analytics workloads, and it inherits the performance engineering that made MinIO popular. The friction is commercial: pricing is capacity-based and targeted at large deployments, which is precisely why smaller shops felt priced out and went looking elsewhere. If you are running hundreds of terabytes to petabytes of AI data on MinIO today, get an AIStor quote and compare it honestly against a migration project — moving petabytes is never free either.

    Community options for labs

    Three open-source projects absorbed most of the MinIO diaspora. SeaweedFS (Apache 2.0, written in Go) is the closest thing to a drop-in replacement and has real production mileage. Garage (AGPLv3, Rust) targets geo-distributed deployments on modest hardware. RustFS (Apache 2.0, Rust) positions itself as a direct MinIO successor and claims better small-object performance, but it is young — treat the claims as unproven until you benchmark them yourself. Community forks of MinIO itself keep binaries flowing, but their security-patching cadence has no track record.

    The line to hold: none of these belongs in a compliance-facing immutability role without a support contract behind it. When the auditor or the cyber-insurance underwriter asks who patches your backup repository, “a GitHub community” is the wrong answer. Labs, dev/test, CI caches — fine. Production backup targets — no.

    What to do about it

    • Inventory every MinIO endpoint. It hides in CI systems, Kubernetes operators, and inside other products’ firmware. You cannot plan a migration you have not scoped.
    • Classify by role. Compliance and immutability workloads migrate first; internal labs can wait or move to community options.
    • Shortlist by use case, not by feature grid. Veeam-only: Ootbi versus ARTESCA. Mixed multi-petabyte: Cloudian versus ARTESCA versus Ceph RGW if you have the ops muscle. Kubernetes-heavy estates should start with our Kubernetes backup buyer’s guide first, because the backup platform choice constrains the storage choice.
    • Test the S3 dialect with real workloads. Versioning, Object Lock in compliance mode, multipart uploads, presigned URLs, and IAM-style policies are where “S3-compatible” claims break.
    • Set a hard deadline. Do not carry archived software past your next penetration test or cyber-insurance renewal. That is the date an unsupported MinIO becomes a written finding.

    The one-line version for the steering committee: MinIO’s archival turned a free dependency into unfunded risk, and the fix is picking a supported platform per workload before the next audit cycle — not finding the single perfect MinIO clone.

    Frequently asked questions

    Is MinIO still safe to use in production?

    It runs, but it is frozen — no guaranteed security patches, no compatibility updates. For internal, low-stakes workloads that is tolerable short-term. For backup repositories, compliance data, or anything internet-adjacent, plan a migration within two to four quarters.

    What is the best alternative to MinIO for enterprise backup?

    For Veeam-only environments, Object First Ootbi is the fastest supported path to an immutable repository. For a general-purpose hardened target that also serves other backup software, Scality ARTESCA is the strongest like-for-like candidate. Cloudian HyperStore wins when one platform must carry backup plus application and AI workloads at scale.

    Is MinIO AIStor free?

    No. AIStor is a commercial, capacity-licensed product aimed at large AI and data-lake deployments. There is no supported free tier equivalent to the old community edition as of mid-2026.

    Can Ceph replace MinIO?

    Yes — Ceph’s RADOS Gateway provides mature S3 compatibility including Object Lock, and commercial support exists via IBM/Red Hat. The cost is operational complexity: Ceph rewards teams with dedicated storage engineering and punishes those without it.

    Does Veeam require Object Lock for immutable backups to object storage?

    Yes. Veeam’s immutability for S3-compatible repositories depends on the target supporting S3 Object Lock with versioning, and Veeam maintains a list of validated object-storage targets. Verify your candidate platform is on it before you buy.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.

  • Is Proxmox Enterprise-Ready? A 2026 Assessment

    Is Proxmox Enterprise-Ready? A 2026 Assessment

    Since Broadcom closed the VMware acquisition, every renewal cycle has produced the same hallway question: could we just run Proxmox? For most of the last two years the honest answer was no — not because the hypervisor was weak, but because the backup row of the RFP matrix killed the evaluation every time. That blocker fell. Veeam Backup & Replication now officially supports Proxmox VE with agentless VM backup and instant restore, and Proxmox VE 9.x shipped the SAN snapshot and SDN capabilities enterprises kept listing as dealbreakers.

    So is Proxmox enterprise-ready in 2026? For specific tiers of workload, yes — and saying otherwise is vendor FUD. For others, not yet — and saying otherwise is homelab enthusiasm wearing a tie. Here is the scorecard, with the thresholds we would use in a real platform decision.

    What changed

    Three things moved between mid-2024 and mid-2026, and together they change the answer.

    • Veeam arrived. Proxmox VE support is built into Veeam Backup & Replication 12.2 and later, and it carried into v13. As of mid-2026 the support matrix covers Proxmox VE 8.2 through 9.1 — agentless backups, changed-block tracking, instant restore, and cross-platform restore into and out of other hypervisors. The single biggest enterprise objection is gone.
    • Proxmox VE 9 shipped grown-up features. Version 9.0 (August 2025, on Debian 13) and 9.1 (November 2025) added snapshot support on thick-provisioned LVM shared storage — the Fibre Channel and iSCSI SAN scenario that used to be an ugly compromise — plus SDN fabrics for scalable network architectures and HA resource affinity rules.
    • Broadcom kept squeezing. Per-core subscription bundling has continued, and renewal quotes at multiples of prior spend are now the norm rather than the horror story. The full field of exits is covered in our VMware alternatives briefing; Proxmox is the one that costs the least to evaluate.

    Why it matters

    Backup support was never a feature-list line item — it was a veto. A hypervisor your data protection vendor will not certify is a hypervisor your auditors, your cyber-insurance carrier, and your DR runbook cannot accept. With Veeam on board, the Proxmox conversation shifts from “we can’t” to “should we” — an operations, support, and skills question rather than a binary disqualifier.

    The market noticed. Proxmox evaluation traffic has climbed steadily since the Broadcom repricing waves, and it is no longer just the sub-100-VM crowd asking. The takeaway for a VP: the question is no longer whether Proxmox can be backed up properly. It is whether your team can run it, and for which tier of workload.

    The 2026 scorecard

    Proxmox VE 9.xvSphere / VCF
    Core hypervisorKVM — mature, performant, effectively at parity for general x86 workloadsESXi — mature, the reference standard
    Cluster schedulingHA plus basic rebalancing; affinity rules new in 9.xDRS — two decades of refinement, still clearly ahead
    Networking at scaleSDN fabrics are new and improving; no true distributed-switch equivalentvDS and NSX — deep, certified, operationally proven
    Shared SAN storageLVM snapshot chains new in 9.x — first-generationVMFS — boring in the best way
    Backup ecosystemVeeam (8.2–9.1) plus Proxmox Backup ServerEvery major vendor, every release, day one
    ISV certificationsThin — few formal support statementsDeep — SAP, Oracle, and the rest of the matrix
    Enterprise supportSubscription tiers, ticket-based, plus partnersGlobal follow-the-sun organization — at a price
    Licensing costPer-socket subscription, low hundreds of euros per socket per year (list pricing varies)Per-core subscription; typically 5–15x more per host

    Where Proxmox genuinely holds up

    Our rule of thumb: if a site runs fewer than roughly 20 hosts per cluster, the workloads are general-purpose, and the team has real Linux depth, Proxmox is a defensible tier-1 choice in 2026 — not a compromise.

    • SMB and mid-market general virtualization. Windows and Linux server VMs, file, print, app servers — KVM handles these at parity, and the license delta funds a lot of engineering time.
    • Edge and ROBO. Two-to-five node sites are where per-core VMware pricing hurts most and where Proxmox with local ZFS or small Ceph shines. This is the easiest first migration.
    • Secondary and DR sites. Veeam’s cross-platform restore makes an asymmetric strategy practical: vSphere in production, Proxmox as the recovery target.
    • Dev/test and internal labs. No serious argument against it — this has been true for years.

    Where it still trails vSphere

    Honest downsides, because the coverage is worthless without them.

    Operations at scale. There is no genuine equivalent of the distributed switch, and lifecycle management across hundreds of hosts — image-based updates, host profiles, staged remediation — is where vSphere’s two-decade head start shows. Proxmox SDN fabrics are promising and first-generation; treat them accordingly.

    The certification matrix. If your ERP or database vendor’s support statement names vSphere and stops there, that is the end of the discussion for that workload. This is not Proxmox’s fault, but it is your problem, and it moves slowly.

    Support depth. Proxmox Server Solutions in Vienna sells subscription tiers with response-time targets, and the partner network is growing — but it is not a global mission-critical support organization, and pretending the two are equivalent is how CIOs get burned. Enterprises that want a full commercial stack with one throat to choke should also price the other exits; our Nutanix versus VMware cost analysis runs those numbers.

    People. The hiring pool of experienced vSphere administrators is enormous. The pool of admins who can debug a Debian-based Ceph cluster at 3 a.m. is not. Budget for training or a support partner — this is the real cost line item most Proxmox TCO decks omit.

    The Veeam reality check

    Veeam’s Proxmox support is real, but read the fine print before you architect around it. The integration runs through a plug-in and worker appliances, and as of mid-2026 Proxmox is managed from the Windows-based backup console — it has not yet reached the web UI of the new Veeam Software Appliance. More important is version lag: Proxmox ships point releases faster than Veeam updates its support matrix, and the gap between a fresh PVE release and official Veeam support is a live topic on both vendors’ forums. The support matrix currently tops out at PVE 9.1.

    The operational rule is the same discipline you learned with vSphere update releases: the backup vendor’s compatibility matrix gates your hypervisor upgrades, not the other way around. Freeze PVE upgrades until Veeam catches up. Proxmox Backup Server remains a solid, deduplicating second option — but a single-vendor backup stack for a single-vendor hypervisor is a concentration risk worth naming out loud.

    What to do about it

    Segment, pilot, and use the leverage — in that order.

    • Segment the estate. Certification-bound and latency-critical tier-1 stays on vSphere for now. Tiers 2 and 3, edge sites, DR targets, and dev/test go on the Proxmox candidate list.
    • Run a 90-day pilot with production-shaped workloads, and make Veeam backup and restore drills — including an instant-restore test under load — a pass/fail gate, not an afterthought.
    • Use the leverage either way. A documented, credible Proxmox pilot changes your Broadcom renewal math even if you never migrate a single production VM. Procurement teams know when a threat is real.
    • Mind the clock. With vSphere 8 support ending in 2027, the platform decision has a deadline. Teams that start evaluating in 2026 decide on their own schedule; teams that wait decide on Broadcom’s.

    Frequently asked questions

    Is Proxmox VE really free for production use?

    The software is open source under the AGPLv3 and free to run. For production, buy a subscription: it provides the enterprise repository — the tested, slower-moving update channel — and access to Proxmox support. Running the no-subscription repository in production is a choice we would not defend in an audit.

    Does Veeam support Proxmox VE?

    Yes. Support is integrated into Veeam Backup & Replication 12.2 and later, including v13, covering Proxmox VE 8.2 through 9.1 as of mid-2026 with agentless backup, changed-block tracking, and instant restore. Always check the current compatibility matrix before upgrading either product.

    Can Proxmox replace VMware in a large enterprise?

    Partially. It is a credible replacement for edge sites, secondary data centers, DR targets, and tier-2/3 workloads today. Wholesale replacement of a large vSphere estate — with its certification, networking, and operations dependencies — is still a multi-year program that most enterprises should not attempt in one step.

    What is the biggest risk of moving to Proxmox?

    Not the hypervisor — KVM is proven at hyperscale. The risk is operational: support depth, upgrade discipline against the Veeam compatibility matrix, and whether your team has (or can hire) the Linux skills to run it at 3 a.m. Price that honestly and the rest of the decision gets much easier.

    Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.