Velero Kubernetes Backup and DR: Velero 1.18 vs Kasten vs CloudCasa
Most teams discover what their Velero Kubernetes backup setup actually protects on the worst possible day: the day they try to restore it. A green “Completed” status tells you that Velero finished writing objects to a bucket. It does not tell you that the persistent volumes came back consistent, that the restore ordering satisfied your admission webhooks, or that you can meet the recovery time your business signed up for. The gap between “backups exist” and “recovery works” is where Kubernetes disaster recovery projects fail.
This matters in 2026 because clusters now carry stateful workloads that used to live on managed databases, including edge and industrial clusters where a rebuild means a truck roll. Velero 1.18 added concurrent backup processing and cache volumes for data movers, and a late patch release fixed a stuck backup queue. Commercial alternatives, Kasten K10 and CloudCasa, compete on the operational layer that open-source Velero leaves to you.
You will leave with a precise mental model of how Velero moves data, a runbook-grade restore procedure, an RPO and RTO planning method, and a decision matrix for Velero versus Kasten versus CloudCasa.
What this covers: Velero architecture and data paths, Kopia and CSI data movement, backup storage locations, hooks, restore ordering, etcd versus Velero scope, runnable CLI and Schedule YAML, restore testing, failure modes, and a three-way tool comparison.
Context and Background
Kubernetes stores its desired state in etcd and its application state wherever your volumes live. Those are two separate recovery problems, and most backup confusion comes from treating them as one. etcd holds the API objects: Deployments, Services, Secrets, CRDs. Persistent volumes hold the bytes your databases and queues wrote. A cluster rebuilt from Git plus a good volume restore is usually a faster path than a cluster rebuilt from an etcd snapshot, and we will return to why.
Velero, an open-source project that began at Heptio and is now maintained as an open-source project on GitHub, is the default answer for the application-level layer. It exports Kubernetes resources through the API server into JSON files, packs them into a tarball in object storage, and coordinates volume data protection through CSI snapshots, a data mover, or file system backup. Its design choice is deliberate: it talks to the Kubernetes API rather than to etcd, so a Velero backup is portable across clusters and Kubernetes versions within the limits of API compatibility.
The commercial field has two notable names. Kasten K10, now sold as Veeam Kasten after Veeam acquired Kasten in 2020, is a policy-driven, application-centric platform with a web UI and database integrations. Its documentation describes backup, restore, disaster recovery and mobility of Kubernetes applications, and lists a free Starter edition alongside paid Enterprise editions. At the time of writing the documentation site shows version 9.0.7 as current. CloudCasa, from Catalogic Software, started as a SaaS backup service and added a management layer for Velero, so teams can keep Velero as the data plane while using CloudCasa as the control plane.
Release status matters for planning. Velero 1.18.0 shipped with the 1.18 blog post dated March 6, 2026, and the 1.18.x line has continued with patch releases. Per the project’s release page, v1.18.4 is the latest patch I could verify, and its only listed change is a fix for a backup queue that became permanently stuck when a dequeued backup completed during a patch operation. That fix is directly relevant to the concurrency work in 1.18, and I will explain why in the failure modes section.
If your clusters run on edge or plant-floor infrastructure, the networking and device layers matter to your recovery design as well. Our write-ups on Cilium versus Calico for industrial Kubernetes networking and on CNI choices across Calico, Cilium, Flannel and Multus cover what you must reproduce on a recovery cluster before any restored pod can talk to anything. For authoritative reference, the Velero documentation is the primary source for every mechanism discussed below.
A note on method. Everything about Velero behavior here comes from the project documentation and release notes I fetched in October 2026. Where I describe commercial products I limit claims to what their public pages state, and I flag anything I could not confirm. Performance figures in this post are either quoted from sources or explicitly illustrative.
How Velero Backs Up a Cluster: The Reference Architecture
A Velero Kubernetes backup is a coordinated operation with three outputs: a tarball of exported Kubernetes resources, volume data captured by one of three mechanisms, and metadata that lets another Velero install discover and restore the backup. The control loop runs in the Velero server pod, the data movement runs in a node-agent daemonset, and everything lands in a BackupStorageLocation, which is an object storage bucket or prefix.

Figure 1: Velero backup flow. The server exports resources and drives hooks and snapshots, node-agent data movers push volume data through Kopia, and both outputs land in the backup storage location.
Figure 1 shows why Velero scales the way it does. Resource export is a set of API calls from one pod, cheap and fast. Volume data is the expensive part, and it is pushed to the node-agent daemonset so that bytes flow from the node that mounts the volume to the bucket, never through the Velero server.
The direct answer: what a Velero backup contains
A Velero backup contains the serialized Kubernetes objects selected by your namespace and label filters, plus volume data captured through a CSI snapshot, a CSI snapshot followed by a data mover upload, or file system backup of mounted pod volumes. It does not contain etcd itself, node configuration, container images, or anything outside the Kubernetes API and the volumes you selected.
That scope statement is the foundation of every design decision that follows. If a thing is not an API object or a selected volume, Velero will not bring it back.
The server, the node-agent, and the custom resources
The Velero server runs controllers for Backup, Restore, Schedule, BackupStorageLocation and related types. When you create a Backup custom resource, directly or through a Schedule, the server lists the resources that match your filters, runs hooks, requests volume snapshots, and writes the result to the bucket. It also runs a periodic sync that reads backup metadata from the bucket and creates Backup objects in the cluster, which is what makes cross-cluster restore possible: install Velero on a new cluster pointing at the same bucket, and the old backups appear.
The node-agent is a Kubernetes daemonset that hosts the data movement controllers. According to the file system backup documentation, it accesses pod volume data through hostPath mounts, typically under /var/lib/kubelet/pods, and launches data mover pods that perform the transfer using Kopia modules. This is why the node-agent needs elevated access on the node, and why hardened clusters that forbid hostPath need explicit exceptions for it.
For CSI snapshot data movement, the documentation describes two custom resources that track the work. A DataUpload is created per CSI snapshot and is watched by data mover controllers on the nodes, ending in Completed, Failed or Cancelled. A DataDownload plays the same role for restores. A BackupRepository resource manages the lifecycle of the repository, and one is created per namespace for CSI snapshot operations. Controllers on different nodes may handle the resource in different phases, but the actual data transfer for one volume runs on a single node.
Backup storage locations and volume snapshot locations
A BackupStorageLocation, or BSL, is described in the Velero documentation as a bucket or a prefix within a bucket under which all Velero data is stored, together with provider-specific fields such as region, plus an optional credential reference. One BSL is the default for the install, and you can change it with the backup-location set command. The BSL status shows Available when Velero has validated access, which you check with the backup-location get command.
A VolumeSnapshotLocation is the sibling concept for providers whose snapshot APIs Velero calls directly, such as cloud disk snapshots. If you use CSI snapshots, the VolumeSnapshotClass in the cluster governs where snapshots live, and the VolumeSnapshotLocation matters far less. Many modern installs never touch it.
The design lesson is to treat the BSL as the most important dependency in your recovery plan. Your backups are only as durable as that bucket, which means versioning, object lock or immutability where the provider offers it, cross-region replication, and credentials that a compromised cluster cannot use to delete history. Velero itself can set a BSL to read-only access mode, which the disaster recovery procedure uses to protect backups during a restore, and we will use that in the runbook.
The encryption caveat you should read before you rely on it
Both the file system backup and CSI data movement documents carry a warning that deserves repeating. Velero’s built-in data movers use a static encryption key across repositories, and the documentation states that anyone who has access to your backup storage can decrypt your backup data. The repository encryption is therefore not a substitute for access control on the bucket.
In practice this pushes you toward bucket-level controls: server-side encryption with customer-managed keys, tight IAM policies, and separate buckets per trust boundary. If your threat model includes a bucket reader who should not see Secrets or database contents, plan for that explicitly. Secrets are exported in the resource tarball as well, so the same reasoning applies to them.
The Three Data Paths for Persistent Volumes
Persistent volumes are where Velero decisions have real cost and risk. Velero 1.18 offers three ways to protect them, and the right choice depends on the storage system, the portability you need, and how fast you must restore.

Figure 2: Choosing a volume data path. Snapshots stay close to the storage system and restore fast, while data movement and file system backup produce portable copies in object storage.
CSI snapshots without data movement
The simplest path asks the CSI driver to take a VolumeSnapshot and leaves it in the storage system. Backup is fast because the snapshot is typically a storage-side metadata operation, and restore is fast for the same reason. The cost is that the snapshot lives with the volume’s storage backend. If the array, the cloud account or the region is lost, the snapshot goes with it, so this path protects against deletion and corruption but not against loss of the storage system.
This mode requires a CSI driver that supports the v1 VolumeSnapshot API and a VolumeSnapshotClass labeled for Velero. It also ties your restore target to the same storage system, which is a limit if your disaster plan involves restoring in a different environment.
CSI snapshot data movement
CSI snapshot data movement, the documentation’s term, combines the snapshot with an upload. Velero creates the CSI snapshot, evaluates volume policies and backup type, creates DataUpload resources, and the data movers transfer the snapshot contents to the backup repository. When the transfer finishes the snapshot is removed. The result is a portable copy in object storage with the consistency point of the snapshot.
The documented prerequisites are Kubernetes 1.20 or later, a CSI driver supporting the v1 VolumeSnapshot API, the MountPropagation feature, a configured BSL on S3-compatible storage, and node-agent installed with the use-node-agent flag. The node-agent-config ConfigMap exposes tuning knobs: priority class for data mover pods, concurrency for parallel uploads and downloads, a limit on the prepare queue, cache volume configuration for restores, and authentication settings for the CSI SnapshotMetadataService.
That last item points to a newer capability. The documentation also describes a block data mover that requires a block storage backend and Linux nodes, with optional support for changed block tracking in CSI. I treat block-level mode as emerging: it exists in the documentation, but I did not verify production maturity across drivers, so test it on your own storage before relying on it.
File system backup
File system backup, abbreviated FSB, reads files from the mounted pod volume rather than from a snapshot. Velero wraps Kopia as a generic file system uploader, and the repository stores deduplicated, encrypted chunks. FSB is storage-agnostic, which makes it the fallback for volumes whose drivers lack snapshot support, and it works across clouds because the data is read through the filesystem.
You select volumes in one of two ways. Opt-in mode requires an annotation on each pod, backup.velero.io/backup-volumes, listing the volumes to include. Opt-out mode backs up all pod volumes by default and excludes the ones named in backup.velero.io/backup-volumes-excludes. Volumes that are excluded or not annotated fall back to snapshot attempts.
The documented limitations matter. FSB cannot back up hostPath volumes, though local volumes are supported. It backs up mounted pod volumes only, so orphaned PVC and PV pairs without a running pod are not captured. Some filesystems, including Azure Files over SMB, blobfuse and gcsfuse, silently lose file ownership on restore. Large files extend the scan needed for deduplication, and the CSI documentation notes that incremental efficiency with the file system data mover is low for large files.
Crash consistency versus application consistency
A snapshot of a running database is crash-consistent: it captures the volume as if the power had been cut. Most modern databases recover from that state, but recovery takes time and some applications will not. Application consistency requires quiescing, which in Velero means backup hooks that freeze or flush before the snapshot and thaw afterward.
Kasten approaches this with blueprints that encode database-specific procedures, and its documentation highlights deep integrations with relational and NoSQL databases. With Velero you build the equivalent yourself using hooks, which we cover next, or you use the database’s own logical backup tooling and treat the volume snapshot as a second line of defense. The honest framing is that Velero gives you the mechanism, and the application knowledge is yours to supply.
What Velero 1.18 changed for data movement
The 1.18 release blog lists several changes relevant to this section. Velero can process multiple backups concurrently, where earlier versions serialized them. Cache volumes can be configured for data mover pods during restore, which addresses pod failures caused by limited ephemeral disk space and concurrency limits on a single node. Incremental size is now reported for both CSI data movement and FSB, so you can see how much data each incremental backup actually moved.
The release also adds glob patterns for namespace filters, so a filter such as team-* can target a family of namespaces, and VolumePolicy supports actions by PVC phase, which lets you skip Pending or Lost PVCs rather than failing a backup on them. On the stability side, the release notes describe running repository operations outside the Velero server process to prevent out-of-memory kills and caching PVC-to-pod lookups to speed VolumePolicy evaluation. The release lists Kopia 0.22.3 and Go 1.25.7. One deprecation to note: the PVC selected-node feature is deprecated, and Velero now handles those annotations automatically.
Velero Kubernetes Backup Runbook: Install, Schedule, Hook, Restore, Verify
This section is the operational core. Everything here is runnable, but treat object names, bucket names and versions as placeholders to adapt, and check flags against velero install --help for your exact version.
Install with node-agent and a bucket
The install command points Velero at a bucket, a provider plugin and a credentials file, and enables the node-agent for data movement. The CSI feature is part of the main Velero server in current releases, so you enable data movement behavior through the node-agent and backup flags rather than a separate plugin in most setups. Confirm this against the 1.18 install documentation for your provider.
velero install \
--provider aws \
--plugins velero/velero-plugin-for-aws:<plugin-version> \
--bucket my-velero-backups \
--backup-location-config region=eu-west-1,s3ForcePathStyle=false \
--secret-file ./credentials-velero \
--use-node-agent \
--default-volumes-to-fs-backup=false \
--features=EnableCSI
velero backup-location get
The final command should show the BSL in the Available phase. If it shows Unavailable, fix credentials and bucket policy before going any further, because every later step depends on it.
For S3-compatible stores such as MinIO or Ceph, add the S3 URL and path-style options to the location config. Those keys are provider plugin settings, and the Velero locations page points to the plugin documentation for the full list rather than enumerating them, so check the plugin README for your store.
A Schedule that matches an RPO
A Schedule is a cron expression plus a Backup template. The example below backs up production namespaces every hour with CSI data movement, keeps backups for seven days, and uses a label selector to exclude noisy namespaces. The hourly cadence is an example, not a recommendation.
apiVersion: velero.io/v1
kind: Schedule
metadata:
name: prod-hourly
namespace: velero
spec:
schedule: "0 * * * *"
useOwnerReferencesInBackup: false
template:
ttl: 168h0m0s
includedNamespaces:
- "prod-*"
excludedResources:
- events
snapshotVolumes: true
snapshotMoveData: true
storageLocation: default
defaultVolumesToFsBackup: false
hooks:
resources:
- name: pg-quiesce
includedNamespaces: ["prod-db"]
labelSelector:
matchLabels:
app: postgres
pre:
- exec:
container: postgres
command: ["/bin/bash", "-c", "psql -U postgres -c 'CHECKPOINT;'"]
onError: Fail
timeout: 60s
Three details here deserve attention. The glob pattern in includedNamespaces relies on the wildcard support added in 1.18, so on older versions you must enumerate namespaces. The snapshotMoveData field is what turns a CSI snapshot into a snapshot plus upload. And the TTL decides when Velero garbage-collects backups, which is your retention policy and also your storage bill.
The default TTL for a Velero backup is 30 days according to the disaster recovery documentation, which refers to daily backups with a default 30-day retention period. Set it explicitly so that nobody has to remember the default.
Hooks for application consistency
Hooks run commands inside pod containers around the backup. The documentation describes annotation-based hooks on pods, with pre hooks running before custom action processing and post hooks after all custom actions complete. The annotations cover the command, the container (defaulting to the first container), on-error behavior (Fail or Continue, defaulting to Fail), and a timeout that defaults to 30 seconds. Hooks can alternatively be declared in the Backup spec, as in the Schedule above, which keeps the logic in version control rather than on workloads.
Commands do not run in a shell by default, so wrap multi-step commands in a bash invocation. Hook results appear in the backup describe output, which reports attempt counts and failures. That output is where you debug a hook that silently timed out.
A pattern that works for databases: a pre hook that flushes or checkpoints, the snapshot, and a post hook that resumes writes. For PostgreSQL the CHECKPOINT shown above reduces recovery time after a crash-consistent restore but does not make the snapshot application-consistent by itself. For true quiescing, use the database’s backup-mode API or a fsfreeze on the data volume, and keep the freeze window short, because the snapshot is requested inside that window.
Run a backup on demand, then inspect it
Before trusting the Schedule, run one backup by hand and read everything it tells you.
velero backup create smoke-$(date +%Y%m%d%H%M) \
--include-namespaces prod-db \
--snapshot-move-data \
--wait
velero backup describe smoke-<timestamp> --details
velero backup logs smoke-<timestamp> | tail -n 50
kubectl -n velero get datauploads
The describe output with details lists the volumes captured and by which mechanism, hook outcomes, and warnings. A backup that finishes as PartiallyFailed usually still contains most resources, but you must read why. The DataUpload listing shows the data movement phase per volume, which is where slow or failed uploads show first.
The restore, in the right order
Velero restores resources in a default order: Custom Resource Definitions, Namespaces, StorageClasses, VolumeSnapshots, PersistentVolumes, PersistentVolumeClaims, RBAC objects, Secrets, ConfigMaps, and finally Pods and Services. You can override this with the restore-resource-priorities flag, where high-priority resources restore first, low-priority last, and unlisted resources sort alphabetically in between.

Figure 3: The recovery sequence on a replacement cluster. Install Velero against the same bucket, protect the backups, restore in dependency order, test, then return to normal operation.
The documented disaster procedure has five steps: schedule regular backups, respond to a disaster, secure the backup storage by setting it to read-only, restore from the most recent backup, and return the storage location to read-write. The read-only patch looks like this.
kubectl patch backupstoragelocation default --namespace velero \
--type merge --patch '{"spec":{"accessMode":"ReadOnly"}}'
velero backup get
velero restore create --from-backup prod-hourly-<timestamp> \
--existing-resource-policy=none \
--namespace-mappings prod-db:prod-db-restored \
--wait
velero restore describe <restore-name> --details
kubectl patch backupstoragelocation default --namespace velero \
--type merge --patch '{"spec":{"accessMode":"ReadWrite"}}'
Read-only mode protects the backup set from being altered or garbage-collected while you work, which matters when a cluster is half-recovered and a stray Schedule could write a bad backup over your good history. Note that a read-only BSL also stops new scheduled backups from writing, so plan the window.
Velero’s default behavior is non-destructive. The restore reference says it never overwrites data that already exists in the cluster: existing resources are skipped, except ServiceAccounts, which are merged. The existing-resource-policy set to update patches existing resources to match the backup, but the documentation calls it best-effort and warns it may not work well for PVCs or Pods. For disaster recovery onto an empty cluster the default is what you want, and for in-place repair you should rehearse the update policy before you need it.
Namespace mapping restores into a different namespace, and PV claim references are updated to match. If a PV already exists on the target cluster and the namespace is remapped, Velero renames it with a velero-clone prefix and a UUID. That behavior is what makes side-by-side restore tests safe, and it is the technique I recommend for restore drills.
Restore hooks run after pod creation and Velero waits for them to finish before continuing. Use them for post-restore initialization, such as running a migration check or warming a cache, and keep them idempotent.
Dependencies Velero will not restore for you
A restored workload needs a cluster that can host it. Before pods can start, the replacement cluster needs the same StorageClass names, the CNI and network policies that your workloads assume, admission controllers and their webhooks, ingress controllers, certificate issuers, and node labels or taints used for scheduling. Velero restores objects in the API, so a Deployment restores without its StorageClass provisioner existing, and then sits Pending.
This is the strongest argument for GitOps as the first layer of recovery. Cluster add-ons, CNI configuration and platform controllers come back from Git through your reconciliation tool, and Velero then restores the namespaces that hold application state. Our coverage of Cilium and Calico choices for industrial clusters explains why network policy is the dependency most often forgotten, and the Azure IoT Akri resource interface post is a good example of device-facing custom resources that must exist before dependent workloads restore.
etcd Snapshots Versus Velero: Two Different Recovery Tools
Teams often ask whether Velero replaces etcd snapshots. It does not, and the answer comes from scope and failure mode, not from preference.
An etcd snapshot captures the entire state of the control plane’s datastore at one moment. Restoring it returns the cluster to that exact state: every object, every resource version, including leases and node registrations. It is the right tool when the control plane itself is destroyed or corrupted, and on managed Kubernetes services you often cannot take one at all because the provider owns etcd.
A Velero backup captures selected API objects as exported documents plus volume data. Restoring it creates objects through the API server, so admission control, defaulting and validation all apply, and you can restore a subset, remap namespaces, or restore into a different cluster version. It cannot restore control plane internals, and it will not bring back cluster-scoped infrastructure that you excluded.
| Dimension | etcd snapshot | Velero backup |
|---|---|---|
| Scope | Entire control plane datastore | Selected namespaces, resources and volumes |
| Includes volume data | No | Yes, by snapshot, data mover or FSB |
| Restore target | Same cluster identity, usually same version | Any compatible cluster |
| Granularity | All or nothing | Namespace, label or resource level |
| Available on managed control planes | Often not | Yes, uses the API |
| Passes admission and validation | No, raw state | Yes, goes through the API |
| Best for | Control plane loss | Application recovery, migration, rebuild |
The practical architecture uses both where you own the control plane: etcd snapshots for control plane loss, Velero for application and volume recovery, and Git for platform configuration. Where you do not own the control plane, Velero and Git carry the whole load, which makes your restore drills more important, not less.
Planning RPO and RTO With Numbers You Can Defend
Recovery point objective is how much data you can lose, and recovery time objective is how long you can be down. For Velero, RPO is bounded by schedule frequency plus backup duration, and RTO is the sum of several stages that you should measure separately.
For RPO, a backup that starts at the top of each hour and takes 12 minutes to finish gives you a recovery point that is, at worst, nearly an hour and 12 minutes old, because the newest complete backup might be the one that finished 59 minutes ago, and a failure just before a backup completes loses the whole window. These numbers are illustrative arithmetic, not benchmarks. The point is that RPO is schedule interval plus backup duration, and any tuning that lengthens the backup lengthens your effective RPO.
For RTO, decompose the recovery into stages and time each one in a drill: provisioning or locating a replacement cluster, installing platform add-ons from Git, installing Velero and syncing backup metadata, restoring resources, restoring volume data, application warm-up and verification, and DNS or traffic cutover. Volume data usually dominates. With CSI snapshots in place the volume restore can be near instant, while with data movement the restore must download from object storage, and its duration scales with data size and available bandwidth.
A working method follows. First, list your workloads and assign each a tier with an RPO and RTO target. Second, map tiers to mechanisms: snapshot-only for tier-three, data movement for tier-two, data movement plus application hooks plus a more frequent schedule for tier-one. Third, drill each tier and record the measured stage durations. Fourth, compare measured RTO with target and adjust, for example by raising data mover concurrency, adding cache volumes for restores, or moving large volumes to storage-native replication.

Figure 4: A decision flow for choosing between Velero, Kasten K10 and CloudCasa based on budget, UI and application-aware needs, and cluster count.
Replication is a legitimate alternative for the tightest objectives. If you need an RPO measured in seconds, backup tools are the wrong layer, and storage-level replication or database-native replication is the right one. Velero and its peers are best understood as protection against deletion, corruption, ransomware and rebuilds, with tight-RPO failover handled elsewhere.
Restore testing as a scheduled job
An untested backup is a hypothesis. Schedule restore tests the way you schedule backups: weekly or monthly for important tiers, automated where possible, and with a pass criterion that is more than “restore completed”.
A solid test restores the latest backup into a scratch namespace using namespace mapping, waits for pods to reach Ready, runs a smoke test against the restored service, and records the elapsed time per stage. The pass criterion should include a data check, for example a row count or checksum compared with production at the backup timestamp. Then it deletes the scratch namespace and emits the timings as metrics so you can chart RTO drift over time.
Do a full cluster rebuild drill at least once or twice a year, because that exercises the dependencies Velero does not cover: GitOps bootstrap, credentials, DNS, and the human runbook. Teams that only test namespace restores discover missing platform pieces during a real incident.
Velero vs Kasten K10 vs CloudCasa: Where Each Fits
The three tools overlap less than the marketing suggests. Velero is a data plane and a set of CRDs. Kasten is a full platform with its own policy engine and interface. CloudCasa can act as a managed control plane over Velero or as a standalone SaaS agent. The right comparison is by operating model, because that determines who does the work.
Operating model
With plain Velero you own everything: installation, upgrades, plugin versions, schedules, alerting, restore drills and cross-cluster inventory. The benefit is transparency and no licensing. The cost is engineering time, and the time grows with cluster count because each cluster has its own Velero install and its own failure modes.
Kasten K10 bundles discovery, policies, a dashboard, application-aware capture and reporting. Its documentation describes automatic application discovery, a policy-driven and extensible architecture, a native Kubernetes API and a web interface, and the vendor page lists a free edition for five nodes. I could not confirm the exact feature split between the free and paid editions from public pages, so check the current edition comparison before you commit to a design that depends on one feature.
CloudCasa for Velero positions itself as a management layer for existing Velero installs, so you keep the open-source engine and add centralized visibility and operations. The public pricing page listed, at the time I fetched it, a free tier for up to five nodes with 30-day retention and a limit of three users, a Pro SaaS tier at 69 US dollars per node per month on annual billing, and a self-hosted Enterprise tier with custom pricing. Treat those prices as a snapshot from a single fetch and confirm them with the vendor, because per-node pricing and tier boundaries change.
Decision matrix
| Criterion | Velero OSS | Kasten K10 | CloudCasa |
|---|---|---|---|
| License model | Open source, no fee | Free edition for five nodes per vendor page, paid tiers beyond | Free tier for five nodes, paid Pro per node, self-hosted Enterprise |
| Interface | CLI and CRDs | Web dashboard and API | SaaS console and API |
| Multi-cluster view | Build your own | Provided by the platform | Core selling point |
| Application-aware capture | Hooks you write | Blueprints and database integrations per vendor docs | Application-aware backups listed on Pro |
| Data engine | Kopia based data movers and CSI snapshots | Its own engine, details in vendor docs | Own agent, or manages Velero |
| Operational burden | Highest | Medium | Lowest for multi-cluster |
| Lock-in | Lowest, open formats and CRDs | Higher, proprietary policy objects | Medium, depends on mode |
| Best fit | Platform teams with Kubernetes depth | Enterprises needing policy, reporting and support | Fleets of clusters wanting central control |
Read the matrix as a starting hypothesis, not a verdict. Vendor-listed capabilities are claims until you test them with your own workloads and storage.
When Velero alone is enough
Velero alone suits teams with a small number of clusters, strong Kubernetes skills, GitOps for platform state and a willingness to write hooks and drills. In that setting the open data formats and absence of per-node pricing are real advantages. It also suits edge deployments with intermittent connectivity where a SaaS control plane is awkward and a local bucket works better.
When to pay for Kasten
Kasten earns its cost when the pain is organizational: dozens of application teams, audit requirements, a need for policy-driven protection that developers cannot misconfigure, and database-specific consistency that you do not want to build yourself. The free five-node edition is a reasonable way to evaluate it on a non-production cluster.
When CloudCasa makes sense
CloudCasa fits when you already run Velero, or want to, across many clusters and your constraint is visibility and operations. It reduces the per-cluster toil of monitoring schedules and failures, and it adds cross-cluster restore workflows. The tradeoff is dependence on a vendor service, so confirm what happens to restores if the SaaS is unreachable during a disaster. That question belongs in your vendor review for any managed control plane.
Trade-offs, Gotchas, and What Goes Wrong
The most common Velero Kubernetes backup failure is a backup that restores objects but not the application. Causes include missing StorageClasses on the target, absent CRDs for operators that were excluded, admission webhooks that reject restored objects before their backing service exists, and Secrets that reference external systems. Each one is a dependency outside Velero’s scope, and each appears only during a real restore or a good drill.
Data mover resource limits are the second trap. Data mover pods run on nodes, consume CPU, memory and ephemeral storage, and compete with workloads. The 1.18 notes mention that restore pods could fail from limited ephemeral disk, which is why cache volumes were added. Set priority classes and concurrency in the node-agent-config ConfigMap deliberately, and watch DataUpload durations as a leading indicator.
Third, the stuck queue. Concurrent backup processing is new in 1.18, and the 1.18.4 patch fixed a case where the backup queue could become permanently stuck when a dequeued backup completed during a patch. The release page lists only that fix, so I cannot describe the triggering conditions in more detail. The operational lesson is to alert on backups that stay in New or InProgress beyond a threshold, and to run the latest patch of your minor version rather than the first release of it.
Fourth, file system backup edge cases. Static encryption keys, silent ownership loss on some network and FUSE filesystems, poor incremental efficiency on large files, and the inability to capture volumes without a mounted pod are all documented. If you store large database files, prefer snapshots or data movement from a snapshot, and test restores of ownership-sensitive data.
Fifth, retention and cost. Backups accumulate: each Schedule run creates a backup that lives until its TTL expires. Object storage costs scale with retained data, request counts and cross-region replication. Velero’s Kopia repositories deduplicate, but the bill depends on churn, not on the logical size of your volumes, and I have no benchmark to quote that would hold across workloads.
Sixth, credentials and blast radius. A cluster that can write backups can often delete them unless the bucket enforces immutability or the credentials lack delete permission. Ransomware playbooks target backup storage first. Use object lock or equivalent, separate accounts, and treat the read-only BSL setting as a recovery aid, not a security control.
Finally, version skew. Restoring a backup taken from a newer Kubernetes API surface onto an older cluster fails for resources that do not exist there, and CRD versions must match what the restored custom resources expect. Keep a note of the API versions in each backup and rehearse cross-version restores if your DR cluster runs a different release.
Practical Recommendations
Start your Velero Kubernetes backup design by separating the three layers: platform state from Git, control plane state from etcd snapshots where you own it, and application state from Velero. Write down which layer restores which object, and make the document part of your runbook. Most surprises come from objects that nobody assigned to a layer.
Choose a Velero Kubernetes backup mechanism per workload tier rather than globally. Use CSI snapshot data movement as the default for stateful workloads that need portability, keep snapshot-only for volumes where deletion protection is enough, and reserve file system backup for storage without CSI snapshot support. Add hooks for databases, and prefer the database’s own consistent-backup mechanism for the most critical data.
Protect the bucket as seriously as production. Enable versioning and object lock where available, restrict delete permissions, replicate to a second region or account, and remember that the built-in repository key is static, so bucket access equals data access. Pin Velero to a current 1.18.x patch, and read the changelog before each upgrade.
Evaluate Kasten and CloudCasa against a written requirement list on a non-production cluster. Both advertise a five-node free tier, so a proof of concept costs little. Measure restore time, application consistency for your database, and behavior when the vendor control plane is unreachable.
A short checklist:
- Every backup location is Available, versioned, and replicated, with delete-protected history.
- Schedules have explicit TTLs and match tier RPOs including backup duration.
- node-agent concurrency, priority class and cache volumes are set and documented.
- A restore into a scratch namespace runs on a schedule with data verification and timing metrics.
- A full cluster rebuild drill runs at least yearly, from an empty cluster, using only the runbook.
- Alerts exist for stuck, partially failed and missing backups.
- Platform add-ons, CNI and StorageClasses are reproducible from Git.
Frequently Asked Questions
What is Velero Kubernetes backup and what does it protect?
Velero Kubernetes backup is done by an open-source tool that backs up and restores Kubernetes resources and persistent volume data. It exports API objects selected by namespace, label or resource type into a tarball in object storage, and captures volumes through CSI snapshots, a Kopia-based data mover, or file system backup. It does not back up etcd, node configuration or container images, so it complements rather than replaces control plane snapshots and GitOps.
Is Velero a replacement for etcd backups?
No. etcd snapshots capture the entire control plane datastore and restore a cluster to an exact prior state, which is the right tool when the control plane is lost. Velero captures selected API objects and volumes and restores them through the API server, which suits application recovery, migration and rebuilds. On managed Kubernetes where you cannot reach etcd, Velero and Git together carry the recovery load.
How does Velero back up persistent volumes?
Velero offers three paths. CSI snapshots stay in the storage system and restore quickly but are not portable. CSI snapshot data movement takes a snapshot and uploads its contents to object storage with Kopia, producing a portable copy. File system backup reads mounted pod volumes directly and works without snapshot support. The node-agent daemonset runs the data movers, and DataUpload and DataDownload resources track each volume transfer.
What changed in Velero 1.18?
Velero 1.18 can process multiple backups concurrently, supports cache volumes for data mover pods during restore, reports incremental backup size for data movers, accepts glob patterns in namespace filters, and lets VolumePolicy act by PVC phase. It also moves repository operations out of the server process to avoid out-of-memory kills. The v1.18.4 patch fixed a backup queue that could become permanently stuck.
Should I choose Velero, Kasten K10 or CloudCasa?
Choose Velero when you have few clusters, strong Kubernetes skills and want open formats without licensing cost. Choose Kasten when policy, reporting, application-aware capture and vendor support matter more than cost. Choose CloudCasa when you run many clusters, want central control over Velero, and accept a SaaS dependency. All three offer a small free tier or trial per public pages, so test restores on your own workloads first.
How do I test a Velero restore without affecting production?
Restore the latest backup into a scratch namespace with the namespace-mappings flag, which remaps PV claim references and renames conflicting PVs with a velero-clone prefix. Wait for pods to become Ready, run a smoke test, compare a data checksum or row count against production at the backup time, record timings per stage, and delete the scratch namespace. Automate it weekly or monthly, and rebuild a full cluster once or twice a year.
Further Reading
- Cilium vs Calico for industrial Kubernetes networking, ADR 2026
- Cilium vs Calico industrial Kubernetes networking, decision record companion
- CNI comparison: Calico, Cilium, Flannel and Multus on Kubernetes
- Azure IoT Akri: a Kubernetes resource interface for devices
- Velero documentation: file system backup, CSI snapshot data movement, restore reference
- Velero 1.18 release blog
- Velero releases on GitHub
By Riju — about
