Skip to content Skip to footer

Velero Backup Automation via Kubernetes CronJob

Velero Backup Automation via Kubernetes CronJob
Velero Backup Automation via Kubernetes CronJob

1. Overview: Why CronJobs Over Native Schedules

Velero includes native backup schedules, but they come with hard limitations: fixed naming patterns, limited operational control, and minimal extensibility when you need to integrate with external systems or customize behavior.

When you need deterministic naming, manual triggerability, or tighter operational control, Kubernetes CronJobs provide a superior abstraction layer.

Key Insight

This article presents a production-ready approach to automating Velero backups using a Kubernetes CronJob without maintaining a custom container image. The setup downloads the Velero CLI dynamically, executes backup commands using a minimal kubectl image, and relies on Kubernetes RBAC for security and isolation.

In order to set up Velero backups we will follow this approach:

  • Downloads the Velero CLI dynamically (v1.13.2) using an init container
  • Executes backup commands using a minimal kubectl image
  • Relies on Kubernetes RBAC for security and isolation
  • Eliminates image maintenance overhead

The result is simple, auditable, and easy to modify—exactly what you want for critical backup infrastructure.

2. Architecture

The workflow is intentionally boring—which is exactly what you want for backups. Complexity is the enemy of reliability.

Kubernetes CronJob
Schedule: 0 0 * * * (Daily at midnight)
Triggers: Job → Pod Creation
↓
Pod: velero-backup
Init Container (fetch-velero)
• Image: alpine:3.20
• Downloads Velero v1.13.2
• Installs to shared volume
• Exit: Signal main container
Main Container (kubectl)
• Image: bitnami/kubectl:latest
• Mounts Velero binary
• Executes: velero backup create
• Uses: velero-backup-sa ServiceAccount
↓
Velero Server
Orchestrates backup workflow
Stores backup data in S3/MinIO
Manages backup lifecycle

✓ Architecture Benefits

  • No sidecars: Simplified pod architecture
  • No custom controllers: Leverage native Kubernetes constructs
  • No image rebuilds: Dynamic CLI fetching eliminates maintenance
  • Clear separation: Init container handles setup, main container executes

3. Key Components: Building Blocks

3.1 ServiceAccount: Identity and Isolation

The CronJob pod needs an identity inside the cluster. That identity is a dedicated ServiceAccount—not default, and definitely not cluster-admin.

kubectl create serviceaccount velero-backup-sa -n velero

This ServiceAccount is used exclusively for backup execution. This isolation principle ensures that if the backup job is compromised, the blast radius is limited to backup operations only.

Security Principle

Never reuse ServiceAccounts across different operational contexts. Each automation should have its own identity with minimal, explicit permissions.

3.2 RBAC: ClusterRole and ClusterRoleBinding

The Velero CLI interacts with cluster-scoped resources: backups, volumes, namespaces, and workloads. RBAC must be explicit and minimal—grant only what's necessary for backup operations.

apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: velero-backup-role rules: - apiGroups: ["velero.io"] resources: - backups - backupstoragelocations - deletebackuprequests - schedules verbs: ["create", "get", "list", "watch"] - apiGroups: [""] resources: - namespaces - pods - persistentvolumes - persistentvolumeclaims verbs: ["get", "list"] - apiGroups: ["apps"] resources: - deployments - statefulsets - daemonsets verbs: ["get", "list"] --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRoleBinding metadata: name: velero-backup-binding roleRef: apiGroup: rbac.authorization.k8s.io kind: ClusterRole name: velero-backup-role subjects: - kind: ServiceAccount name: velero-backup-sa namespace: velero

Apply and verify permissions:

kubectl apply -f sa-role.yaml # Verify ServiceAccount can list deployments kubectl auth can-i list deployments \ --as=system:serviceaccount:velero:velero-backup-sa

Critical Checkpoint

If the permission check fails, do not proceed. Debug RBAC before deploying the CronJob. A backup job that silently fails due to permission issues is worse than no automation at all.

RBAC Component Purpose Scope
ClusterRole Defines permissions for Velero operations Cluster-wide resources
ClusterRoleBinding Binds permissions to ServiceAccount Links role to identity
ServiceAccount Pod identity for backup execution Namespace: velero

3.3 Init Container: Dynamic Velero CLI Fetching

The init container solves a common problem cleanly: the main image doesn't ship with the Velero binary, and you don't want to maintain a custom image.

Using alpine:3.20 as the base, the init container:

  1. 1 Installs minimal dependencies (curl and tar)
  2. 2 Downloads Velero v1.13.2 from official GitHub releases
  3. 3 Extracts and installs the binary into a shared emptyDir volume
  4. 4 Exits cleanly, signaling the main container to start

This approach keeps the main container minimal and avoids custom image builds, reducing your operational surface area.

initContainers: - name: fetch-velero image: alpine:3.20 command: ["/bin/sh", "-c"] args: - | set -eux apk add --no-cache curl tar curl -L -o /tmp/velero.tar.gz \ https://github.com/vmware-tanzu/velero/releases/download/v1.13.2/velero-v1.13.2-linux-amd64.tar.gz tar -xzf /tmp/velero.tar.gz -C /tmp install -m 0755 /tmp/velero-v1.13.2-linux-amd64/velero /target/velero volumeMounts: - name: velero-bin mountPath: /target

Version Pinning

The Velero version (v1.13.2) is explicitly pinned in the download URL. When upgrading, update this version string and test thoroughly in a non-production environment first.

3.4 Main Container: Executing the Backup

The main container uses bitnami/kubectl, mounts the shared volume containing the Velero binary, and executes the backup command. The container design is intentionally simple—one responsibility, clearly defined.

Namespace-scoped backup example:

BACKUP_NAME="polaris-backup-$(date -I)" velero backup create "$BACKUP_NAME" \ --include-namespaces polaris \ --default-volumes-to-fs-backup \ --wait

Cluster-wide backup example (same CronJob pattern, different command):

BACKUP_NAME="cluster-backup-$(date -I)" velero backup create "$BACKUP_NAME" \ --include-cluster-resources=true \ --default-volumes-to-fs-backup \ --wait

Nothing is hardcoded except intent. The backup name includes a date stamp for easy identification, and the --wait flag ensures the job doesn't exit until the backup completes or fails.

Velero Flag Purpose Use Case
--include-namespaces Limit backup scope to specific namespaces Application-specific backups
--include-cluster-resources Include cluster-scoped resources (CRDs, PVs) Full cluster disaster recovery
--default-volumes-to-fs-backup Use file-system backup for PVs (not snapshots) Cloud-agnostic volume backups
--wait Block until backup completes Ensures job reflects actual backup status

4. Complete CronJob Manifest

Here's the full, production-ready manifest that ties everything together:

apiVersion: batch/v1 kind: CronJob metadata: name: velero-backup-polaris namespace: velero spec: schedule: "0 0 * * *" successfulJobsHistoryLimit: 2 failedJobsHistoryLimit: 2 jobTemplate: spec: template: spec: serviceAccountName: velero-backup-sa restartPolicy: Never volumes: - name: velero-bin emptyDir: {} initContainers: - name: fetch-velero image: alpine:3.20 command: ["/bin/sh", "-c"] args: - | set -eux apk add --no-cache curl tar curl -L -o /tmp/velero.tar.gz \ https://github.com/vmware-tanzu/velero/releases/download/v1.13.2/velero-v1.13.2-linux-amd64.tar.gz tar -xzf /tmp/velero.tar.gz -C /tmp install -m 0755 /tmp/velero-v1.13.2-linux-amd64/velero /target/velero volumeMounts: - name: velero-bin mountPath: /target containers: - name: kubectl image: bitnami/kubectl:latest command: ["/bin/sh", "-c"] args: - | set -eux BACKUP_NAME="polaris-backup-$(date -I)" velero backup create "$BACKUP_NAME" \ --include-namespaces polaris \ --default-volumes-to-fs-backup \ --wait volumeMounts: - name: velero-bin mountPath: /usr/local/bin

✓ Manifest Highlights

  • Schedule: Cron expression for daily midnight execution
  • History Limits: Keeps 2 successful and 2 failed jobs for debugging
  • Restart Policy: Never—failed backups should be investigated, not auto-retried
  • Volume Sharing: emptyDir enables binary sharing between init and main containers

5. Operational Procedures

5.1 Schedule Configuration

The default schedule runs daily at midnight UTC (0 0 * * *). Adjust based on your backup requirements and maintenance windows.

Schedule Cron Expression Use Case
Daily at midnight 0 0 * * * Standard daily backups
Every 6 hours 0 */6 * * * High-frequency production environments
Every 12 hours 0 */12 * * * Balanced backup frequency
Weekly (Sunday midnight) 0 0 * * 0 Less critical environments

List active CronJobs:

kubectl get cronjobs -n velero

5.2 Logs and Troubleshooting

When backups fail (and they will), you need quick access to diagnostic information.

View recent jobs:

kubectl get jobs -n velero

Check main container logs (backup execution):

kubectl logs job/<job-name> -n velero -c kubectl

Check init container logs (Velero download):

kubectl logs job/<job-name> -n velero -c fetch-velero

Silent Failures

If Velero fails silently (exits 0 despite backup failure), your automation is lying to you. Always verify backup completion through the Velero API, not just pod exit codes.

Verify backup completion:

velero backup get velero backup describe <backup-name>

5.3 Cleanup Policy

The manifest includes sensible history limits to prevent cluster clutter while retaining enough data for debugging:

successfulJobsHistoryLimit: 2 failedJobsHistoryLimit: 2

This configuration keeps the two most recent successful jobs and the two most recent failed jobs. Adjust based on your debugging workflow—more history means more kubectl noise.

6. Security Considerations

Security is not an afterthought—it's baked into the architecture from the start.

Security Layer Implementation Why It Matters
Dedicated ServiceAccount velero-backup-sa (not default) Limits blast radius if compromised
Minimal RBAC Only backup-related permissions Prevents privilege escalation
No cluster-admin Explicit resource access only Principle of least privilege
Namespace Isolation Runs in velero namespace Separation from application workloads

Security Anti-Pattern

If your backup job needs cluster-admin, your security model is fundamentally broken. Backup operations should never require unrestricted cluster access. If you find yourself granting cluster-admin, stop and redesign your RBAC strategy.

Additional Security Hardening

  • Image Pinning: Use digest-based image references instead of :latest tags
  • Network Policies: Restrict pod network access to only Velero API endpoints
  • Pod Security Standards: Apply restricted pod security admission
  • Audit Logging: Enable Kubernetes audit logs for backup job activity

7. Testing: Trust but Verify

Automation without testing is just scheduled failure. Always validate your backup pipeline before relying on it in production.

7.1 Manual Test Execution

Create an ad-hoc job from the CronJob template to test immediately without waiting for the schedule:

kubectl create job \ --from=cronjob/velero-backup-polaris \ manual-test-backup -n velero

Monitor job execution:

kubectl get jobs -n velero -w kubectl logs job/manual-test-backup -n velero -c kubectl -f

7.2 Verify Backup Creation

After the job completes, confirm the backup was created and is valid:

velero backup get velero backup describe polaris-backup-$(date -I)

Check backup status fields:

  • Phase: Should be "Completed"
  • Errors: Should be 0
  • Warnings: Investigate any warnings
  • Total Items: Should match expected resource count

7.3 Test Restore Operations

Critical Reality Check

If you don't test restores, this entire setup is theater. A backup system that can't restore is worse than no backup system at all—it provides false confidence.

Perform a test restore to a separate namespace:

velero restore create test-restore-$(date +%s) \ --from-backup polaris-backup-$(date -I) \ --namespace-mappings polaris:polaris-restore

Verify restored resources:

kubectl get all -n polaris-restore velero restore describe test-restore-$(date +%s)
Test Type Frequency Purpose
Backup Creation After every CronJob change Verify job executes successfully
Restore to Separate Namespace Monthly Validate backup integrity
Full Disaster Recovery Drill Quarterly Test complete recovery procedures
RBAC Validation After permission changes Ensure ServiceAccount has required access

8. Advanced Patterns and Customization

8.1 Multiple Backup Schedules

Deploy multiple CronJobs with different schedules and scopes for layered backup strategy:

  • Namespace-specific: Frequent backups (every 6 hours) for critical apps
  • Cluster-wide: Daily full cluster backups
  • Selective resources: Hourly backups of specific resource types

8.2 Notification Integration

Extend the main container to send notifications on backup completion or failure:

# Add to main container args if velero backup create "$BACKUP_NAME" --wait; then curl -X POST "$SLACK_WEBHOOK" \ -d "{\"text\":\"✓ Backup $BACKUP_NAME completed successfully\"}" else curl -X POST "$SLACK_WEBHOOK" \ -d "{\"text\":\"✗ Backup $BACKUP_NAME FAILED - investigate immediately\"}" exit 1 fi

8.3 Backup Retention Automation

Add a second CronJob to clean up old backups automatically:

# Run weekly to delete backups older than 30 days velero backup delete --all --older-than 720h

9. Troubleshooting Common Issues

Symptom Cause Solution
Init container fails to download Velero Network policy blocking GitHub access Allow egress to github.com in NetworkPolicy
Permission denied errors RBAC misconfiguration Verify ClusterRoleBinding with kubectl auth can-i
Backup created but empty Velero server configuration issues Check Velero server logs and BackupStorageLocation status
Job succeeds but no backup visible Velero CLI version mismatch Ensure CLI version matches server version
CronJob never triggers Invalid cron syntax Validate cron expression with online tools

10. Conclusion: Simple, Auditable, Maintainable

This approach to Velero backup automation delivers on three critical principles:

  1. Simplicity: No custom images, no complex controllers, just standard Kubernetes constructs
  2. Auditability: Every component is visible, every permission is explicit
  3. Maintainability: Updating Velero versions is a one-line change, no image rebuilds required

Production Readiness Checklist

  • ✓ ServiceAccount created with dedicated identity
  • ✓ RBAC permissions tested and verified minimal
  • ✓ CronJob schedule aligned with backup requirements
  • ✓ Manual test job executed successfully
  • ✓ Backup creation verified in Velero
  • ✓ Restore tested to separate namespace
  • ✓ Monitoring and alerting configured
  • ✓ Disaster recovery procedures documented

The beauty of this solution is in what it doesn't require: no custom container builds, no image registry management, no version tracking headaches. When Velero releases a new version, you update a single URL in the init container and test. That's it.

For production Kubernetes environments, backup automation isn't optional—it's foundational. This CronJob-based approach provides that foundation without introducing unnecessary complexity or operational overhead.

Remember: Untested backups are aspirational, not operational. Build the automation, test the restores, and sleep better knowing your disaster recovery strategy is more than documentation.

Leave a Comment