This is the security hardening playbook for multi-tenant Kubernetes platforms. Not a checklist from a compliance slide deck. These are the defense layers, policy decisions, and architectural patterns that keep production platforms secure when dozens of teams share the same cluster — and attackers get faster every quarter.
In the previous post, we covered scaling a Kubernetes platform to 20,000+ users — control plane hardening, multi-tenancy isolation, autoscaling, and observability. But scaling without security is just building a bigger target.
Here's the uncomfortable truth: Kubernetes is not secure by default. Pods run as root. Containers can access the host network. There's no admission control unless you install it. Network traffic between namespaces flows freely. And every new namespace you onboard expands your attack surface.
This post covers the six security layers that every multi-tenant Kubernetes platform needs — from admission control through runtime detection. We'll walk through the architecture decisions, explain why each layer matters, and be honest about what the new Kubernetes-native security features actually deliver versus what still requires external tooling.
1. Defense-in-Depth — The Only Security Model That Works
Security at scale isn't one tool or one policy — it's layered defense where each layer catches what the previous one missed. This is defense-in-depth applied to Kubernetes, and it's non-negotiable for any platform serving multiple teams.
The four layers work in sequence. Layer 1 (Network Edge) handles perimeter security — firewall rules, WAF, and TLS termination at the ingress controller. Layer 2 (Admission Control) is the gatekeeper — every resource request passes through the API server, where Kyverno policies and Pod Security Admission validate it against your security baseline before anything is created. Layer 3 (Workload Isolation) enforces boundaries between tenants — NetworkPolicies, ResourceQuotas, and RBAC prevent lateral movement and resource abuse. Layer 4 (Runtime Security) is the last line — Falco with eBPF monitors syscalls in real-time, detecting container escapes, unexpected process execution, and anomalous network connections that bypass every other layer.
Each layer is independent. If an attacker gets past your ingress controls, admission policies still block privileged containers. If a misconfigured policy lets a bad workload through, runtime detection catches the shell spawn. No single layer is sufficient — but together, they make exploitation significantly harder and detection significantly faster.
The 90-Second Rule
Research consistently shows that new Kubernetes clusters face their first automated attack attempt within minutes of exposure. If your security posture requires manual intervention to enforce, it's already too late. Every layer described in this post must be automated — deployed via GitOps, enforced at admission, and monitored continuously.
2. Pod Security Standards — The Baseline Everyone Skips
When Kubernetes removed Pod Security Policies (PSPs) in v1.25, it replaced them with Pod Security Admission (PSA) — a simpler, built-in admission controller that enforces three predefined security profiles: Privileged, Baseline, and Restricted.
Most teams know PSA exists. Almost nobody enforces it properly.
The Three Profiles and When to Use Each
Privileged is unrestricted — pods can do anything including accessing the host network, running as root, and mounting host paths. This profile should only exist in your monitoring namespace (for DaemonSets like Falco that need host access) and nowhere else.
Baseline blocks the most dangerous configurations — no privileged containers, no host networking, no host PID/IPC sharing. It's the minimum acceptable profile for any namespace, but it still allows running as root inside the container, which is a significant gap.
Restricted is what production namespaces should enforce. It requires runAsNonRoot: true, drops all capabilities, sets allowPrivilegeEscalation: false, enforces a read-only root filesystem, and mandates seccomp profiles. This is the profile that actually prevents the container escape exploits that make headlines.
| Profile | Allows Root? | Allows Host Network? | Allows Privilege Escalation? | Use For |
|---|---|---|---|---|
| Privileged | Yes | Yes | Yes | System DaemonSets only (Falco, node exporters) |
| Baseline | Yes | No | Yes | Legacy apps during migration — temporary only |
| Restricted | No | No | No | All production tenant namespaces |
Why PSA Alone Isn't Enough
Pod Security Admission only validates pods — not Deployments, not ConfigMaps, not Ingress resources. It can't enforce image registry allowlists, require specific labels, or generate resources automatically. It's a floor, not a ceiling. For a multi-tenant platform, you need PSA as the baseline plus a policy engine like Kyverno for everything PSA can't cover.
Enforcement Modes Matter
PSA supports three modes per namespace: enforce (reject violations), audit (log violations but allow), and warn (show warnings to the user). The recommended rollout for production is to start with warn on Restricted, identify which workloads fail, fix them, move to audit to catch remaining issues in logs, and only then flip to enforce. Jumping straight to enforcement on a running cluster will break workloads — guaranteed.
3. Policy-as-Code with Kyverno — The Admission Gatekeeper
PSA handles pod-level security profiles. Kyverno handles everything else — and in a multi-tenant platform, "everything else" is where the real risk lives.
The choice between Kyverno and OPA/Gatekeeper comes down to operational reality. OPA requires learning Rego, a domain-specific query language. Kyverno policies are written in YAML — the same language your team already uses for every Kubernetes resource. For platform teams that need to ship security policies quickly and have them maintained by the same engineers who write Helm charts, Kyverno wins on day-two operations every time.
The Five Policies Every Multi-Tenant Platform Needs
These are the non-negotiable admission policies. If your platform doesn't enforce all five from day one, you have a security gap that will be found — either by an auditor or by an attacker.
1. Block Privileged Containers — No container should ever run with privileged: true unless explicitly exempted for system workloads. This single policy prevents the most common container escape path. Kyverno can enforce this with a validation rule that matches on Pod resources and denies any security context with privileged set to true.
2. Enforce Image Registry Allowlists — Only images from your approved registries (Harbor, ECR, GCR) should be admitted to the cluster. This prevents developers from pulling random images from Docker Hub that may contain malware or unpatched vulnerabilities. The policy validates that every container image reference starts with your approved registry prefix.
3. Require Resource Requests and Limits — Every container must declare CPU and memory requests and limits. Without this, a single pod can consume all available resources on a node, causing evictions across tenant namespaces. This policy catches what LimitRange defaults might miss.
4. Generate Default NetworkPolicies on Namespace Creation — This is where Kyverno's generate capability shines. Whenever a new namespace is created, Kyverno automatically generates a default-deny NetworkPolicy. No human remembers to add this on every namespace. Automation makes the secure path the default path.
5. Verify Image Signatures with Cosign — Only images signed by your CI pipeline should be deployed. Kyverno's verifyImages rule checks Cosign signatures at admission time, ensuring that unsigned or tampered images are rejected before they ever run. This is supply chain security enforced at the cluster edge.
Kyverno vs. OPA/Gatekeeper — The Honest Comparison
Use Kyverno if your team is Kubernetes-native and wants YAML policies with validation, mutation, and resource generation in one tool. It's simpler to adopt, easier to maintain, and covers 90% of policy use cases. Use OPA/Gatekeeper if you need complex cross-resource logic, external data integration, or if your organization already has a Rego policy library. Use both if you have the team capacity — Kyverno for operational policies, OPA for compliance-grade logic. Most platform teams start with Kyverno and never need Gatekeeper.
4. PodCertificateRequests — Native mTLS Without a Service Mesh
This is the feature that changes the game for multi-tenant security in 2026. PodCertificateRequests, introduced as beta in Kubernetes v1.35 (KEP-4317), provide a built-in mechanism for pods to obtain TLS certificates directly from the Kubernetes control plane — no Istio, no Linkerd, no external PKI infrastructure required.
How It Works
The flow is elegant in its simplicity. When a pod requests a certificate, the kubelet generates a key pair on the node, creates a PodCertificateRequest resource via the API server, and an external signer controller (you choose the CA) processes the CSR and returns a signed certificate. The kubelet then writes the certificate bundle directly to the pod's filesystem. The API server enforces node restriction at admission time, ensuring that a compromised node can only request certificates for pods running on that node.
The result: pods can authenticate to the API server using mTLS instead of service account tokens, and pods can potentially authenticate to each other using the same certificates. No sidecar proxies. No control plane overhead from a service mesh. Just certificates delivered by the Kubernetes machinery that's already running.
What It Actually Replaces (And What It Doesn't)
Let me be direct about what PodCertificateRequests deliver today and what they don't.
What it delivers: A native, secure mechanism for pod-to-API-server mTLS authentication. This eliminates bearer token replay attacks for workloads that access the Kubernetes API. It also provides the infrastructure for external projects to build certificate distribution on top of Kubernetes primitives instead of rolling their own.
What it doesn't replace: A full service mesh. PodCertificateRequests don't give you transparent pod-to-pod mTLS, traffic management, observability, or retry logic. If you need automatic mTLS between every microservice without application changes, you still need Istio or Linkerd. The KEP explicitly lists pod-to-pod mTLS as a non-goal for the initial implementation.
The Honest Assessment
PodCertificateRequests are a foundational primitive, not a complete solution. They're valuable for workloads that need to authenticate to the API server securely, and they lay the groundwork for future Kubernetes-native service mesh alternatives. But if someone tells you that PodCertificateRequests replace Istio today, they haven't read the KEP. The ecosystem around this feature is still nascent — the signer controller options are limited and mostly demo-grade. Give it 12–18 months before relying on it for production mTLS between services.
Where PodCertificateRequests Shine Today
The immediate value is in eliminating service account token exposure. In multi-tenant platforms, service account tokens are long-lived bearer tokens that can be stolen and replayed. With PodCertificateRequests, workloads authenticate to the API server via mTLS — the private key never leaves the node, and the certificate has a bounded lifetime. For environments with strict compliance requirements (finance, healthcare, government), this is a meaningful security improvement available today.
5. Runtime Security — Detecting What Admission Control Misses
Admission control is preventive. It stops bad configurations from entering the cluster. But it can't stop a legitimate workload from being compromised after deployment — a vulnerable dependency exploited via an API, a supply chain attack that passes signature verification, or a zero-day in the container runtime itself.
Runtime security is the detection layer. It watches what's actually happening inside containers and alerts when behavior deviates from expected patterns.
Falco + eBPF — Syscall-Level Visibility
Falco runs as a DaemonSet on every node, using eBPF probes to intercept syscalls from every container. It doesn't modify container behavior — it observes and alerts. The detection rules are declarative: define what normal looks like, and everything else triggers an alert.
The detections that matter most for multi-tenant platforms:
Shell spawned in container — Production containers should never spawn a shell. If /bin/sh or /bin/bash executes inside a running container, it's either an attacker or someone doing kubectl exec in production (which should also trigger an alert and a conversation).
Unexpected outbound network connections — Your API container should connect to your database and your cache. If it suddenly connects to an IP address in a country you don't operate in, something is wrong. Falco can detect new network connections that don't match expected patterns.
File writes to sensitive paths — Containers with read-only root filesystems shouldn't write to /etc, /usr/bin, or /root. Any write to these paths indicates either a misconfiguration or an active exploit installing persistence.
Privilege escalation attempts — Processes attempting to change UID, load kernel modules, or access /proc/sysrq-trigger are attempting container escape. Falco detects these at the syscall level before they succeed.
The eBPF Advantage
Older versions of Falco used a kernel module for syscall interception, which required kernel headers and introduced stability risks. Modern Falco uses eBPF — no kernel module required, no risk of crashing the node, and significantly lower overhead. On production nodes handling tenant workloads, eBPF-based Falco adds less than 2% CPU overhead. That's a remarkably low cost for syscall-level visibility across your entire platform.
From Detection to Response
Detection without response is just noise. Falco alerts should flow into Alertmanager, which triggers PagerDuty for critical detections (shell spawn, privilege escalation) and creates Slack notifications for informational ones (unexpected network connections). For mature platform teams, automated remediation — killing the compromised pod, cordoning the node, or scaling down the affected deployment via ArgoCD — turns detection into containment in seconds rather than hours.
6. Supply Chain Security — Trust Nothing You Didn't Build
The most sophisticated admission control in the world is useless if the container image itself is compromised. Supply chain attacks — injecting malicious code into base images, dependencies, or build pipelines — have become the most effective attack vector against Kubernetes platforms.
The Three Pillars of Supply Chain Security
Vulnerability Scanning (Trivy) — Every image must be scanned for known CVEs before it enters your registry. Trivy integrates into CI pipelines and scans both OS packages and application dependencies. The critical decision is your threshold: block on Critical and High CVEs, allow Medium with a remediation deadline, and accept Low as informational. Scanning must happen at build time (in CI) and continuously (in the registry) because new CVEs are published daily against images that were clean yesterday.
Image Signing (Cosign) — After an image passes scanning, your CI pipeline signs it with Cosign using either keyless signing (via Fulcio + OIDC identity) or a static key pair. The signature is stored alongside the image in your OCI registry. This creates a cryptographic chain of trust: only images that were built by your pipeline and passed your quality gates carry a valid signature.
Signature Verification at Admission (Kyverno) — The final enforcement point. Kyverno's verifyImages policy checks every image reference at admission time. No valid Cosign signature? The pod is rejected. This closes the loop: an attacker who compromises a developer's credentials and pushes a malicious image directly to the registry can't deploy it — because it wasn't signed by the CI pipeline.
The SBOM Requirement Is Coming
Software Bill of Materials (SBOMs) — machine-readable inventories of every component in your container images — are transitioning from "nice to have" to compliance requirement. The US Executive Order on Cybersecurity, the EU Cyber Resilience Act, and FedRAMP all reference SBOM requirements. Generate SBOMs with Syft at build time, store them in your registry alongside the image, and have a plan for querying them when the next Log4Shell-class vulnerability drops and leadership asks "are we affected?" within 30 minutes.
Harbor as the Hardened Registry
Public registries (Docker Hub, GitHub Container Registry) are convenience tools, not security boundaries. A multi-tenant platform needs a private registry with vulnerability scanning, signature storage, RBAC per project, replication policies, and audit logging. Harbor delivers all of this and runs on Kubernetes itself. Every tenant team gets their own project with scoped access — no team can push to or pull from another team's project.
The registry is the chokepoint where supply chain security either holds or fails. If your registry accepts unsigned images, your entire signing infrastructure is theater.
Security Hardening Checklist
| # | Practice | Priority | Layer |
|---|---|---|---|
| 1 | Pod Security Admission — Restricted profile on all tenant namespaces | 🔴 Critical | Admission |
| 2 | Kyverno — block privileged containers, enforce registry allowlists | 🔴 Critical | Admission |
| 3 | Default-deny NetworkPolicies auto-generated per namespace | 🔴 Critical | Isolation |
| 4 | Image scanning with Trivy — block Critical/High CVEs | 🔴 Critical | Supply Chain |
| 5 | Image signing with Cosign — verified at admission by Kyverno | 🟠 High | Supply Chain |
| 6 | Falco with eBPF — runtime syscall monitoring on all nodes | 🟠 High | Runtime |
| 7 | RBAC audit — no cluster-admin outside platform team | 🟠 High | Access Control |
| 8 | PodCertificateRequests — mTLS for API server authentication | 🟡 Medium | Identity |
| 9 | SBOM generation with Syft — stored in registry | 🟡 Medium | Supply Chain |
| 10 | Harbor private registry — per-team project isolation | 🟡 Medium | Supply Chain |
The Security Maturity Curve
Not every platform needs every layer from day one. Security maturity is progressive. Here's the recommended order of implementation:
Month 1 (Foundation): Pod Security Admission in Restricted mode. Default-deny NetworkPolicies. RBAC audit — remove unnecessary cluster-admin bindings. These three changes cost almost nothing and close the biggest gaps.
Month 2–3 (Admission Control): Deploy Kyverno with the five core policies. Set up Trivy scanning in your CI pipeline. Start signing images with Cosign. This layer prevents bad workloads from entering the cluster.
Month 4–6 (Runtime + Supply Chain): Deploy Falco with eBPF. Configure alerting for shell spawns and privilege escalation. Implement signature verification at admission. Set up Harbor if you're still pulling from public registries. Generate SBOMs.
Month 6+ (Advanced): Evaluate PodCertificateRequests for API server mTLS. Implement automated remediation for runtime detections. Set up continuous compliance scanning with Kyverno's audit mode. Build security dashboards in Grafana.
The Honest Cost
Full security hardening adds approximately 15–20% overhead to your platform engineering workload. Kyverno and Falco each need dedicated configuration and ongoing policy maintenance. Image signing requires CI pipeline changes. And someone needs to respond to the alerts — security tooling without a response process is just expensive logging. Budget for 1 dedicated security-focused platform engineer for every 50+ namespaces.
Wrapping Up
Kubernetes security for multi-tenant platforms is defense-in-depth — layered controls where each layer compensates for the gaps in the previous one. Pod Security Admission sets the baseline. Kyverno enforces organizational policies at admission. NetworkPolicies isolate tenants. Falco detects runtime compromise. And supply chain security ensures the images running in your cluster are the ones your pipeline built and signed.
Start with the foundation: Restricted PSA profiles and default-deny NetworkPolicies. Layer in Kyverno and image scanning. Add runtime detection when you have the team capacity to respond to alerts. And be honest about PodCertificateRequests — they're a promising primitive for API server mTLS, not a service mesh replacement. Yet.
The secure path should always be the default path. If developers have to opt into security, they won't. If security is the only option, you've won.
