Skip to content Skip to footer

Managing IPsec Tunnels with an Operator

Managing IPsec Tunnels with an Operator

Part 1 ended at a wall. The StrongSwan sidecar is a fine answer for one or two tunnels, but every new external peer meant another pod to template, another proxy to reason about, and a lifecycle stitched together by hand. This post is about what you build when that stops being tenable: an operator that treats an IPsec tunnel as a first class Kubernetes object.

When one pod stops being enough

The sidecar works because it hides a tunnel inside a pod. That's also exactly why it doesn't scale. Each new integration is a fresh copy of the same imperative bundle — a StrongSwan config, an Nginx proxy, a probe script, a Deployment — and the only thing holding it together is the person who remembers how it fits. Ten tunnels is ten pods, ten proxy surfaces, ten places a source address gets silently rewritten, and ten little lifecycles nobody reconciles.

The symptoms show up fast. Config drifts between sidecars because someone patched one pod's cipher proposal and forgot the others. Pre-shared keys rotate on different schedules because each rotation is a manual redeploy. Nobody can answer a question as basic as "which external tunnels do we run, and are they up right now?" without grepping Deployments and shelling into pods to read ipsec status. And because every consuming workload reaches the remote service through an Nginx hop, the remote side sees the proxy's address, not the client's — which quietly breaks any allowlist the partner maintains on their end.

The deeper problem underneath all of these is that the sidecar describes a tunnel as a collection of scripts and config, not as a piece of desired state. Nothing in the cluster knows what "the partner-billing tunnel should be up" means. If it drops, a pod restarts and hopes. There is no object to point at, no status to query, no controller whose entire job is to keep the declared truth true.

Many near-identical StrongSwan sidecar pods sprawling across a cluster toward different remote appliances
The sidecar's cost is linear and its state lives in people's heads. Each new peer is another pod to template, key to rotate, and proxy to babysit.

The pattern: tunnels as custom resources

The operator model inverts the sidecar. Instead of packaging a tunnel inside a workload, you declare it. A Custom Resource Definition introduces an IPsecTunnel object, and creating a tunnel becomes a kubectl apply — remote gateway, local identity, IKE proposals, traffic selectors, and a reference to the key material, all as fields on a spec. A controller watches those objects and does whatever it takes to make reality match them, forever.

Here's the shape of it. Everything sensitive stays a reference, not a literal; everything environment-specific is a field, not a rebuild:

# One object describes the whole tunnel. Promote it between environments by # changing values, never code — exactly the discipline from Part 1, now declarative. apiVersion: ipsec.clustercraftops.io/v1alpha1 kind: IPsecTunnel metadata: name: partner-billing spec: remote: gateway: 203.0.113.10 # far-side appliance (WatchGuard/Fortinet/ASA) id: "billing.partner.example" local: id: "@cluster-egress" ike: version: 2 proposals: [aes256-sha256-modp2048] # must match the peer exactly dpd: 30s # dead-peer detection interval auth: pskSecretRef: partner-billing-psk # Secret, never an inline key selectors: - local: 10.90.7.0/24 # managed pool (what workloads source from) remote: 172.16.4.0/24 # the remote service subnet ipam: pool: 10.90.7.0/24

StrongSwan doesn't disappear; it moves. Rather than one charon daemon per sidecar, a single IKE daemon runs centrally under the operator's control, and the controller drives it programmatically over the vici interface — StrongSwan's versatile configuration API. That one detail matters more than it looks: the sidecar approach ultimately meant scraping the text output of ipsec status to guess whether a tunnel was healthy. With vici the controller loads connections, initiates security associations, and subscribes to structured SA-up / SA-down events. The tunnel stops being a container you ship and becomes a record the platform reconciles.

IPsec operator architecture with CRDs, controller, charon over vici, per-connection XFRM interfaces and VXLAN segments
Declare a tunnel as a custom resource; the controller programs the IKE daemon over vici and stitches the data plane — one XFRM interface and one VXLAN segment per connection.

Why an operator instead of just more sidecars

More sidecars scale the work linearly; an operator scales the abstraction. The reconciliation loop that keeps one tunnel healthy is the same loop that keeps fifty healthy. You write the "how" once, inside the controller, and every future tunnel is just another row of desired state — no new code, no new pod template, no new probe script, no new key-rotation runbook.

Anatomy of the operator

Underneath the tidy CRD there are four moving parts, and it's worth being honest about each, because the operator's power is exactly proportional to how much kernel-level plumbing it takes on your behalf.

The controller is the brain. It watches IPsecTunnel resources, runs a reconcile loop against a work queue, and uses finalizers so that deleting a tunnel actually tears down its SA, interface, and overlay instead of leaking them. The IKE daemon (charon) negotiates and maintains the security associations, handles rekeying and dead-peer detection, and is driven entirely over vici so the controller never shells out. Per connection, the operator programs an XFRM interface — a routable Linux netdevice tagged with a unique interface ID (if_id) — so traffic for a given tunnel is selected by which interface it leaves on, not by brittle policy matching on subnets. That interface is a real ip link you can attach routes to and point tcpdump at, which turns debugging from guesswork into inspection. Finally, a VXLAN segment (its own VNI per connection) carries a workload's packets from wherever its pod happens to be scheduled to the node where that tunnel's XFRM interface lives — so any pod in the cluster can use a tunnel without every node needing to terminate IPsec or hold NET_ADMIN.

The reconcile loop itself is the shape every operator has, applied to networking state:

1 · Observe — diff the desired IPsecTunnel spec against charon's actual SA state (via vici)
2 · Program — load/update the connection, bring up the XFRM interface and VXLAN segment, install routes
3 · Verify — confirm the SA is established and the remote selector is reachable end to end
4 · Report — write tunnel health to the resource's status subresource, emit events, then requeue

Every resync runs the same loop, which is what makes drift self-correcting: if someone deletes an XFRM interface by hand, or charon rekeys and the peer misbehaves, the next reconcile notices the gap between desired and actual and closes it — without a human and without a redeploy.

Operator reconcile loop: observe, program, verify, report, then requeue
Observe, program, verify, report — then do it again. The loop that heals one tunnel is the loop that heals all of them, and it runs on every resync whether or not anything changed.

How a workload joins a tunnel

This is the part that makes the operator worth the machinery. In the sidecar world a consuming workload connected to an Nginx ClusterIP and never knew a VPN existed — clean, but it cost you a proxy hop and the original client address. The operator keeps the ignorance and drops the hop. A workload opts into a tunnel with a single annotation:

metadata: annotations: ipsec.clustercraftops.io/tunnel: partner-billing

A mutating admission webhook intercepts the pod at creation, pulls an address from that tunnel's managed pool, and wires the pod onto the tunnel's VXLAN segment. From then on the pod talks to the remote service at its real destination address, from a real, stable source IP the remote side can filter on — no stream proxy in the middle rewriting anything, and nothing in the application aware that a tunnel exists. IP management, which was a manual afterthought in Part 1, becomes a first-class feature: each tunnel owns a pool, the operator hands out addresses as pods start, and a finalizer reclaims them when pods go away, so the pool never silently exhausts from leaked leases.

Pod annotation triggers a mutating webhook, IP allocation from a managed pool, and attachment to a VXLAN segment
Annotate a pod, get an address from the tunnel's pool, join the segment — no proxy hop, real source IPs the remote can allowlist, and IPAM the platform manages for you.

The privilege moved — it didn't shrink

In the sidecar, NET_ADMIN was fenced inside one throwaway pod. The operator centralizes that privilege into a long-lived, cluster-wide control plane, which is a far more valuable target. That's a real trade, and the mitigation is discipline: tight RBAC on the CRDs and the vici socket, NET_ADMIN scoped only to the data-plane pods (not the controller itself), dedicated nodes for tunnel termination, and a mutating webhook you treat as security-critical — with rotated serving certs and a fail-closed policy. An operator you can't audit is worse than the sidecars it replaced.

What makes it production-ready: reconciliation

Part 1's sidecar self-healed the only way a pod can — by dying and coming back. A dropped security association meant the liveness probe failed, Kubernetes restarted the pod, and the tunnel re-established from scratch, taking every connection on it down in the process. That's acceptable for one tunnel. It's a blunt instrument for fifty.

The operator heals with a scalpel. Because the controller continuously compares charon's live SA state against the desired spec, a tunnel that drops is simply re-initiated in place — the offending connection comes back without restarting a pod and without disturbing the other tunnels sharing the daemon. Dead-peer detection and rekey events arrive over vici, so a silently dropped SA is a signal the controller acts on, not a timeout a workload discovers the hard way.

And because every tunnel is a real object, its health is a real field. Wire up status conditions and printer columns and the platform answers the question the sidecar never could:

$ kubectl get ipsectunnel NAME PHASE SA REKEYS REMOTE AGE partner-billing Established 42 3 203.0.113.10 6d legacy-db Established 17 9 198.51.100.7 6d appliance-emea Connecting - - 192.0.2.44 12s

Underneath those columns are status.conditions (Established, Degraded, LastError), Kubernetes Events explaining every transition, and Prometheus metrics falling out of the reconcile loop for free — SA up/down gauges, rekey counters, DPD timeouts, IPAM pool utilisation. You alert on the metric, not on a user ticket.

The payoff

A dropped SA becomes a self-correcting blip on one tunnel instead of a pod restart that ripples across all of them. New integrations are a few lines of YAML, not a new deployment. Key rotation is editing one Secret the controller re-reads. And for the first time you can answer "which tunnels do we have and are they healthy?" with a single command — because the platform, not a person, is holding the state.

Where even the operator strains

The operator is the right answer once integrations multiply, and for many platforms it's the last answer they'll ever need. But it buys that power with concentration and complexity, and it's worth naming the seams before you commit.

  • The data plane centralizes. Traffic funnels through the node(s) terminating the tunnels, which becomes a bandwidth ceiling and a failure domain you now have to think about spreading and scaling. A single SA can only live in one place, so tunnel-node HA is active/passive with a re-establish on failover, not seamless.
  • You own a lot of kernel. XFRM interfaces plus a VXLAN overlay per connection is real plumbing to debug at 2am, and it's version-sensitive — if_id-based XFRM interfaces need a reasonably modern kernel, and the people who can reason about it are rarer than the people who can read a Deployment.
  • IPAM is now your problem at platform scale. Pools overlap, exhaust, and collide with remote expectations; managing address space is a permanent operational surface, not a one-time setup.
  • It's still a VPN bolted onto Kubernetes. However elegant the CRD, the underlying model is "make the cluster act like a branch office," rather than expressing external connectivity in the cluster's own native terms.

The gotcha that will page you: MTU

Stack a VXLAN header (~50 bytes) on top of ESP overhead on top of the underlay, and full-size packets no longer fit. The tunnel comes up, small requests work, then a large POST or a TLS handshake with big certificates hangs forever. Plan MTU end to end from day one — lower the pod MTU on the segment, or clamp TCP MSS — because this failure is intermittent, payload-dependent, and looks like everything except what it is.

That "VPN bolted on" point is the one that matters most, and it's where the series goes next. Instead of teaching Kubernetes to run IPsec, Part 3 asks what external connectivity looks like when you lean on Kubernetes-native constructs — egress gateways, the Gateway API, and mesh-level egress — and when reaching for a general pattern beats hand-owning an IKE data plane at all.

IPsec operator pattern versus Kubernetes-native external connectivity comparison
Two more answers to the same problem. Part 2 is the operator; Part 3 goes Kubernetes-native.
Pattern B · this post

IPsec operator

Tunnels as custom resources, a controller driving charon over vici, XFRM and VXLAN per connection, pool-based IPAM, and workloads that opt in by annotation. Scales when integrations multiply — at the cost of a centralized data plane you own.

Pattern C · next post

Kubernetes-native egress

External connectivity expressed with native primitives — egress gateways, Gateway API, mesh egress — instead of a hand-built IKE data plane. Best when you want to stop owning kernel plumbing and let the platform's own abstractions carry the load.


This is Part 2 of a 3-part series on external connectivity from Kubernetes.
Part 1 — IPsec out of Kubernetes with StrongSwan
Part 2 — Managing IPsec tunnels with an operator (you're reading it)
Part 3 — External connectivity, the Kubernetes-native way (coming next)

Leave a Comment