1. Overview: Why CronJobs Over Native Schedules
Velero includes native backup schedules, but they come with hard limitations: fixed naming patterns, limited operational control, and minimal extensibility when you need to integrate with external systems or customize behavior.
When you need deterministic naming, manual triggerability, or tighter operational control, Kubernetes CronJobs provide a superior abstraction layer.
Key Insight
This article presents a production-ready approach to automating Velero backups using a Kubernetes CronJob without maintaining a custom container image. The setup downloads the Velero CLI dynamically, executes backup commands using a minimal kubectl image, and relies on Kubernetes RBAC for security and isolation.
In order to set up Velero backups we will follow this approach:
- Downloads the Velero CLI dynamically (v1.13.2) using an init container
- Executes backup commands using a minimal
kubectlimage - Relies on Kubernetes RBAC for security and isolation
- Eliminates image maintenance overhead
The result is simple, auditable, and easy to modify—exactly what you want for critical backup infrastructure.
2. Architecture
The workflow is intentionally boring—which is exactly what you want for backups. Complexity is the enemy of reliability.
• Image: alpine:3.20
• Downloads Velero v1.13.2
• Installs to shared volume
• Exit: Signal main container
• Image: bitnami/kubectl:latest
• Mounts Velero binary
• Executes: velero backup create
• Uses: velero-backup-sa ServiceAccount
✓ Architecture Benefits
- No sidecars: Simplified pod architecture
- No custom controllers: Leverage native Kubernetes constructs
- No image rebuilds: Dynamic CLI fetching eliminates maintenance
- Clear separation: Init container handles setup, main container executes
3. Key Components: Building Blocks
3.1 ServiceAccount: Identity and Isolation
The CronJob pod needs an identity inside the cluster. That identity is a dedicated
ServiceAccount—not default, and definitely not cluster-admin.
This ServiceAccount is used exclusively for backup execution. This isolation principle ensures that if the backup job is compromised, the blast radius is limited to backup operations only.
Security Principle
Never reuse ServiceAccounts across different operational contexts. Each automation should have its own identity with minimal, explicit permissions.
3.2 RBAC: ClusterRole and ClusterRoleBinding
The Velero CLI interacts with cluster-scoped resources: backups, volumes, namespaces, and workloads. RBAC must be explicit and minimal—grant only what's necessary for backup operations.
Apply and verify permissions:
Critical Checkpoint
If the permission check fails, do not proceed. Debug RBAC before deploying the CronJob. A backup job that silently fails due to permission issues is worse than no automation at all.
| RBAC Component | Purpose | Scope |
|---|---|---|
| ClusterRole | Defines permissions for Velero operations | Cluster-wide resources |
| ClusterRoleBinding | Binds permissions to ServiceAccount | Links role to identity |
| ServiceAccount | Pod identity for backup execution | Namespace: velero |
3.3 Init Container: Dynamic Velero CLI Fetching
The init container solves a common problem cleanly: the main image doesn't ship with the Velero binary, and you don't want to maintain a custom image.
Using alpine:3.20 as the base, the init container:
- 1 Installs minimal dependencies (
curlandtar) - 2 Downloads Velero v1.13.2 from official GitHub releases
- 3 Extracts and installs the binary into a shared
emptyDirvolume - 4 Exits cleanly, signaling the main container to start
This approach keeps the main container minimal and avoids custom image builds, reducing your operational surface area.
Version Pinning
The Velero version (v1.13.2) is explicitly pinned in the download URL. When upgrading, update this version string and test thoroughly in a non-production environment first.
3.4 Main Container: Executing the Backup
The main container uses bitnami/kubectl, mounts the shared volume containing the Velero binary, and
executes the backup command. The container design is intentionally simple—one responsibility, clearly defined.
Namespace-scoped backup example:
Cluster-wide backup example (same CronJob pattern, different command):
Nothing is hardcoded except intent. The backup name includes a date stamp for easy
identification, and the --wait flag ensures the job doesn't exit until the backup completes or fails.
| Velero Flag | Purpose | Use Case |
|---|---|---|
--include-namespaces |
Limit backup scope to specific namespaces | Application-specific backups |
--include-cluster-resources |
Include cluster-scoped resources (CRDs, PVs) | Full cluster disaster recovery |
--default-volumes-to-fs-backup |
Use file-system backup for PVs (not snapshots) | Cloud-agnostic volume backups |
--wait |
Block until backup completes | Ensures job reflects actual backup status |
4. Complete CronJob Manifest
Here's the full, production-ready manifest that ties everything together:
✓ Manifest Highlights
- Schedule: Cron expression for daily midnight execution
- History Limits: Keeps 2 successful and 2 failed jobs for debugging
- Restart Policy: Never—failed backups should be investigated, not auto-retried
- Volume Sharing: emptyDir enables binary sharing between init and main containers
5. Operational Procedures
5.1 Schedule Configuration
The default schedule runs daily at midnight UTC (0 0 * * *). Adjust based on your backup
requirements and maintenance windows.
| Schedule | Cron Expression | Use Case |
|---|---|---|
| Daily at midnight | 0 0 * * * |
Standard daily backups |
| Every 6 hours | 0 */6 * * * |
High-frequency production environments |
| Every 12 hours | 0 */12 * * * |
Balanced backup frequency |
| Weekly (Sunday midnight) | 0 0 * * 0 |
Less critical environments |
List active CronJobs:
5.2 Logs and Troubleshooting
When backups fail (and they will), you need quick access to diagnostic information.
View recent jobs:
Check main container logs (backup execution):
Check init container logs (Velero download):
Silent Failures
If Velero fails silently (exits 0 despite backup failure), your automation is lying to you. Always verify backup completion through the Velero API, not just pod exit codes.
Verify backup completion:
5.3 Cleanup Policy
The manifest includes sensible history limits to prevent cluster clutter while retaining enough data for debugging:
This configuration keeps the two most recent successful jobs and the two most recent failed jobs. Adjust based on your debugging workflow—more history means more kubectl noise.
6. Security Considerations
Security is not an afterthought—it's baked into the architecture from the start.
| Security Layer | Implementation | Why It Matters |
|---|---|---|
| Dedicated ServiceAccount | velero-backup-sa (not default) | Limits blast radius if compromised |
| Minimal RBAC | Only backup-related permissions | Prevents privilege escalation |
| No cluster-admin | Explicit resource access only | Principle of least privilege |
| Namespace Isolation | Runs in velero namespace | Separation from application workloads |
Security Anti-Pattern
If your backup job needs cluster-admin, your security model is fundamentally broken. Backup operations should never require unrestricted cluster access. If you find yourself granting cluster-admin, stop and redesign your RBAC strategy.
Additional Security Hardening
- Image Pinning: Use digest-based image references instead of
:latesttags - Network Policies: Restrict pod network access to only Velero API endpoints
- Pod Security Standards: Apply restricted pod security admission
- Audit Logging: Enable Kubernetes audit logs for backup job activity
7. Testing: Trust but Verify
Automation without testing is just scheduled failure. Always validate your backup pipeline before relying on it in production.
7.1 Manual Test Execution
Create an ad-hoc job from the CronJob template to test immediately without waiting for the schedule:
Monitor job execution:
7.2 Verify Backup Creation
After the job completes, confirm the backup was created and is valid:
Check backup status fields:
- Phase: Should be "Completed"
- Errors: Should be 0
- Warnings: Investigate any warnings
- Total Items: Should match expected resource count
7.3 Test Restore Operations
Critical Reality Check
If you don't test restores, this entire setup is theater. A backup system that can't restore is worse than no backup system at all—it provides false confidence.
Perform a test restore to a separate namespace:
Verify restored resources:
| Test Type | Frequency | Purpose |
|---|---|---|
| Backup Creation | After every CronJob change | Verify job executes successfully |
| Restore to Separate Namespace | Monthly | Validate backup integrity |
| Full Disaster Recovery Drill | Quarterly | Test complete recovery procedures |
| RBAC Validation | After permission changes | Ensure ServiceAccount has required access |
8. Advanced Patterns and Customization
8.1 Multiple Backup Schedules
Deploy multiple CronJobs with different schedules and scopes for layered backup strategy:
- Namespace-specific: Frequent backups (every 6 hours) for critical apps
- Cluster-wide: Daily full cluster backups
- Selective resources: Hourly backups of specific resource types
8.2 Notification Integration
Extend the main container to send notifications on backup completion or failure:
8.3 Backup Retention Automation
Add a second CronJob to clean up old backups automatically:
9. Troubleshooting Common Issues
| Symptom | Cause | Solution |
|---|---|---|
| Init container fails to download Velero | Network policy blocking GitHub access | Allow egress to github.com in NetworkPolicy |
| Permission denied errors | RBAC misconfiguration | Verify ClusterRoleBinding with kubectl auth can-i |
| Backup created but empty | Velero server configuration issues | Check Velero server logs and BackupStorageLocation status |
| Job succeeds but no backup visible | Velero CLI version mismatch | Ensure CLI version matches server version |
| CronJob never triggers | Invalid cron syntax | Validate cron expression with online tools |
10. Conclusion: Simple, Auditable, Maintainable
This approach to Velero backup automation delivers on three critical principles:
- Simplicity: No custom images, no complex controllers, just standard Kubernetes constructs
- Auditability: Every component is visible, every permission is explicit
- Maintainability: Updating Velero versions is a one-line change, no image rebuilds required
Production Readiness Checklist
- ✓ ServiceAccount created with dedicated identity
- ✓ RBAC permissions tested and verified minimal
- ✓ CronJob schedule aligned with backup requirements
- ✓ Manual test job executed successfully
- ✓ Backup creation verified in Velero
- ✓ Restore tested to separate namespace
- ✓ Monitoring and alerting configured
- ✓ Disaster recovery procedures documented
The beauty of this solution is in what it doesn't require: no custom container builds, no image registry management, no version tracking headaches. When Velero releases a new version, you update a single URL in the init container and test. That's it.
For production Kubernetes environments, backup automation isn't optional—it's foundational. This CronJob-based approach provides that foundation without introducing unnecessary complexity or operational overhead.
Remember: Untested backups are aspirational, not operational. Build the automation, test the restores, and sleep better knowing your disaster recovery strategy is more than documentation.
