Skip to content Skip to footer

IPsec Out of Kubernetes with StrongSwan

IPsec Out of Kubernetes with StrongSwan

Sooner or later a cluster has to talk to something that isn't in the cluster — a billing system behind a partner's firewall, a legacy database that will never get a public endpoint, an appliance that only speaks IPsec. This is the first post in a short series on doing that properly. We start where almost everyone starts: a StrongSwan tunnel running as a sidecar.

The problem nobody designs for up front

Most platform architectures assume the interesting traffic flows into the cluster. Ingress controllers, service meshes, load balancers — the whole toolbox points inward. Then a product team needs to reach a database sitting on a partner network, reachable only across a site-to-site IPsec tunnel terminated on a hardware appliance, and suddenly the cluster has to behave like a branch office on someone else's VPN.

You can't put an Ingress in front of that. The remote side expects encrypted ESP packets and a negotiated IKE security association, not an HTTP request. The traffic has to originate from inside the cluster, get encrypted, traverse the tunnel, and come back — transparently enough that the consuming workload doesn't need to know any of it is happening.

Cluster workloads blocked from reaching a firewalled external service
The shape of the problem: workloads need a private external service that no amount of Ingress will reach.

The pattern: StrongSwan as a dedicated sidecar

The cleanest first answer is to stop trying to make every pod VPN-aware and instead build one small, dedicated pod whose entire job is to be the tunnel. StrongSwan — the long-standing open-source IKEv2 implementation — establishes and maintains the IPsec security association to the remote gateway. A lightweight Nginx stream proxy sits in front of it and exposes the remote service inside the cluster as an ordinary ClusterIP.

That second part is the move that makes the whole thing maintainable. Consuming workloads connect to a normal Kubernetes Service on a normal port. They have no idea a VPN exists. All the privileged, fiddly, kernel-level networking is fenced inside a single pod in its own namespace, and everything else in the cluster stays blissfully ignorant.

StrongSwan plus Nginx sidecar architecture
One pod owns the tunnel. StrongSwan negotiates IPsec; Nginx forwards TCP from a ClusterIP into it.

Why a TCP stream proxy rather than letting pods route directly into the tunnel? Because routing arbitrary pod traffic across an XFRM interface means touching the node's routing tables and handing out NET_ADMIN far more widely than you'd like. Funnelling through Nginx keeps the blast radius to exactly one pod, gives you a clean place to terminate health checks, and means the remote service appears as a single stable address regardless of how the tunnel underneath is behaving.

Why StrongSwan over WireGuard here

WireGuard is faster and simpler when you own both ends. But the appliance on the far side of a corporate tunnel almost always speaks IPsec/IKEv2 and nothing else. StrongSwan exists precisely for the case where the remote peer is non-negotiable — a WatchGuard, a Fortinet, a Cisco ASA — and you have to meet it on its terms.

One image, two environments

Staging and production rarely point at the same remote endpoint, and they often differ in how many external services they expose. The temptation is to build two images. Don't. A single container with a startup script that reads an ENVIRONMENT variable and selects the right tunnel and proxy configuration keeps the two environments honestly identical in everything except the values that genuinely differ.

Single image selecting staging or production config at startup
The same artifact ships to both environments; a startup variable decides which tunnel and which proxy ports come up.

The connection parameters — remote address, local identity, the pre-shared key, the traffic selectors — all arrive as environment values templated in at deploy time. Nothing sensitive lives in the image, and promoting a build from staging to production changes configuration, never code. This is the difference between a VPN pod you trust and one you're afraid to redeploy on a Friday.

The privilege you can't avoid

StrongSwan needs the NET_ADMIN capability to install its routing and security-policy entries. There's no getting around it for this pattern. The mitigation is containment: keep it on this one pod, in its own namespace, with nothing else co-scheduled. A capability granted to one tightly-scoped tunnel pod is a very different risk than one sprinkled across a deployment.

What makes it production-ready: probes

A VPN tunnel is a stateful thing that fails in quiet, annoying ways. The IKE daemon can be alive while the tunnel is down. The remote peer can drop the security association and not tell you. Nginx can be happily listening while there's nothing behind it. A naive liveness check that only asks "is the process running?" will report green while every connection times out.

The fix is a single probe script that returns success only when two independent conditions both hold: the remote target is actually reachable through the tunnel, and the local proxy is serving its health route. Wire that script into Kubernetes startup, readiness, and liveness probes and the platform does the rest — holding traffic until the tunnel is genuinely up, and restarting the pod when it silently dies.

Health probe checking both tunnel reachability and proxy health
A probe that passes only when both the tunnel and the proxy are healthy turns a fragile pod into a self-healing one.

The payoff

With probes wired correctly, a dropped tunnel becomes a non-event: the readiness probe pulls the pod out of rotation, liveness restarts it, the security association re-establishes, and traffic resumes — usually before anyone notices. No pager, no manual ipsec up at 2am.

Where this pattern stops scaling

The sidecar is the right first answer, and for one or two external services it may be the right permanent answer. It's lightweight, it leans entirely on Kubernetes primitives, and there's nothing exotic to learn. But it carries assumptions that strain as the number of tunnels grows.

  • Every new external peer is another pod to template, deploy, and reason about. The Nginx proxy becomes a config surface that grows linearly with your integrations.
  • The proxy is a real hop. Anything that needs the original client IP, or a protocol the stream proxy doesn't cleanly forward, fights the architecture.
  • There's no IP management. If the remote side expects specific source addresses per service, you're hand-assigning them.
  • The lifecycle is yours. Loading config and bringing tunnels up is scripted by hand rather than reconciled by a controller.

This is exactly the point where a Kubernetes-native approach starts to look attractive: describe the tunnel as a custom resource, let a controller run the IKE daemon and program per-connection XFRM interfaces and VXLAN segments, and let workloads opt into a tunnel with a pod annotation that hands them an address from a managed pool. That's the operator model — and it's where this series goes next.

Sidecar pattern versus operator pattern comparison
Two answers to the same problem. Part 1 is the sidecar; Part 2 moves to the operator.
Pattern A · this post

StrongSwan sidecar

One pod, one (or a few) tunnels, an Nginx proxy, and standard probes. Minimal moving parts. Best when you have a small, stable set of external services and want zero new abstractions.

Pattern B · next post

IPsec operator

Tunnels as custom resources, a controller managing XFRM and VXLAN per connection, pool-based IP allocation, and workloads that opt in by annotation. Scales when integrations multiply.

Connecting a cluster to something it wasn't built to reach?

This is the kind of edge that quietly eats weeks of platform time. If you're weighing the sidecar against an operator for your own external-connectivity problem, let's talk it through.


This is Part 1 of a 3-part series on external connectivity from Kubernetes.
Part 1 — IPsec out of Kubernetes with StrongSwan (you're reading it)
Part 2 — Managing IPsec tunnels with an operator (coming next)
Part 3 — External connectivity, the Kubernetes-native way

Leave a Comment