Parts 1 and 2 built increasingly good answers to the same question: how do you get an IPsec tunnel out of a Kubernetes cluster and keep it alive? This post asks a harder one. whether you should have been building a tunnel at all. For a growing share of external dependencies the answer is no, and the cluster already has better primitives for the job.
The question the operator couldn't answer
The operator from Part 2 is genuinely good engineering. Tunnels are declarative, self-healing, observable, and workloads join them with an annotation. If you have a fleet of IPsec peers, build it — nothing in this post retracts that.
But step back and look at what it costs to run. You own an IKE daemon, per-connection XFRM interfaces, a VXLAN overlay, an IPAM system, a mutating webhook, and an MTU problem. That is a small networking product, maintained by a platform team whose actual job is something else. And it's worth asking honestly what fraction of your external dependencies genuinely require it.
For most teams the answer is uncomfortable. The legacy appliance that started this whole series does need IPsec — it speaks nothing else. But the partner API, the payment provider, the object storage endpoint, the observability vendor, the internal service in a peered VPC: those all speak TLS. Routing them through an IKE data plane you hand-built doesn't make them more secure. It makes them dependent on kernel plumbing that has nothing to do with the guarantee anyone actually wanted, which was this identified workload may talk to that specific endpoint, encrypted, and I can prove it happened.
The pattern: make egress a policy, not a tunnel
The Kubernetes-native answer inverts the question. Instead of asking "how do I build a private path to that system," you ask "which workload identity is allowed to reach which destination, over what encryption, and how do I see it." Those are all things the cluster can already express.
Three primitives carry almost all of it. NetworkPolicy gives you default-deny egress, so nothing leaves a namespace unless you said it could. An egress gateway — a dedicated set of proxy replicas on their own nodes — becomes the single funnel through which permitted traffic exits, giving the outside world one stable source address to allowlist. And workload identity, issued as short-lived certificates, means the policy decision is made against who the caller is rather than which IP it happened to get from the CNI.
That last shift is the substantive one. In Part 1, the remote side saw an Nginx proxy's address. In Part 2 it saw an address from a managed pool, which was better. Here it sees a stable egress IP and the call itself carries a cryptographic identity the mesh verified before it left. The IP allowlist stops being your security model and becomes what it should always have been: a coarse network control layered under a real one.
Why this is different from just "put a proxy in front of it"
Part 1's Nginx was also a proxy, and it was a real hop that erased the client. The difference isn't the proxy — it's that the egress gateway is told what to do by declarative objects rather than a config file baked into an image, and that it makes decisions on verified identity rather than on the source IP of whatever pod happened to connect. Same topology, completely different operational model.
Gateway API, pointed outward
Most platform teams already know the Gateway API from ingress. The useful realisation is that the same object model describes egress: a Gateway with a listener, routes that match destinations, and policies attached to backends. You are not learning a new abstraction — you are pointing an existing one in the other direction, which means the same review process, the same GitOps repo, and the same mental model your team already has.
A permitted external dependency becomes a few objects rather than a tunnel:
Compare that to Part 2's IPsecTunnel. Both are declarative, both live in Git, both are reconciled by a controller. The difference is what happens underneath: the operator programmed kernel state you own and debug, while this programs a userspace proxy the platform already runs for ingress. The blast radius of a mistake is a rejected connection, not a broken route table on a node.
Anatomy of a native egress path
It's worth walking a single request end to end, because the interesting work happens in places the application never sees.
Step 5 is the one that quietly justifies the whole exercise. With a tunnel, "who called the partner API and when?" is answerable only if the application logged it. Here it falls out of the data path: request counts and latency per identity and destination, TLS handshake failures, certificate expiry, policy denials. You alert on a metric before the partner emails you, and when someone asks which services depend on a vendor you are about to drop, you query it instead of grepping code.
The payoff
Adding an external dependency becomes a pull request against a route, reviewed like any other change, with no kernel state, no address pool, and no MTU maths. Removing one is deleting an object. And for the first time the question "what does this cluster talk to on the outside, and who is allowed to?" has an answer that lives in the platform rather than in a wiki someone stopped updating in 2024.
Where the native approach doesn't reach
This is the part the vendor diagrams leave out. Native egress is the right default, not a universal replacement, and pretending otherwise is how platform teams end up rebuilding a tunnel badly at the worst possible moment.
- The remote may simply not speak TLS. The appliance from Part 1 wants ESP and IKEv2 and will not be talked out of it. No amount of Gateway API changes that — you still need Pattern A or B for that peer.
- Non-HTTP protocols get awkward. A raw database connection or a legacy binary protocol can be passed through at L4, but you lose most of the routing and observability that made the approach attractive in the first place.
- The remote may demand a fixed private address. Some partners allowlist an RFC1918 source that only exists inside their VPN. A stable public egress range doesn't satisfy that requirement.
- You've traded kernel complexity for control-plane complexity. A mesh or Gateway controller is not free. It's a well-supported product rather than bespoke code, which is a real improvement, but it is still a thing to run, upgrade, and debug.
The failure mode to design against: silent expiry
Native egress leans on certificates — the mesh's own short-lived identities, and the CA bundle you pinned for the remote. Both expire. A partner rotating their CA without telling you produces a total, instant outage for exactly one dependency, and the error surfaces as a TLS handshake failure buried in proxy logs. Alert on remote certificate expiry and on handshake-failure rate per destination from day one. This is the native equivalent of Part 2's MTU trap: rare, sharp, and obvious only in hindsight.
Choosing between the three patterns
Three posts, three answers, and the useful conclusion is not that the last one wins. They're a ladder, and the discipline is picking the lowest rung that actually solves your problem.
| Question | Sidecar (A) | Operator (B) | Native egress (C) |
|---|---|---|---|
| Remote speaks | IPsec only | IPsec only | TLS / HTTP |
| Scales to many peers | No — linear cost | Yes | Yes |
| You own kernel state | Some (one pod) | Yes — xfrm, VXLAN, IPAM | No |
| Identity-aware policy | No | No — network-level only | Yes |
| Per-call observability | Process probes only | Tunnel-level metrics | Per request and identity |
| Sharp edge to plan for | Proxy hides client IP | MTU with VXLAN + ESP | Certificate expiry |
| Reach for it when | 1–2 fixed IPsec peers | A fleet of IPsec peers | Everything else — the default |
The honest end state for most platforms is a mix. Run native egress as the default path for the long tail of TLS dependencies, where it gives you policy, identity and telemetry for far less operational weight. Keep an IPsec capability — a sidecar if it's one or two peers, the operator if it's many — for the appliances that will never change. What you should not do is what the series opened with: build a VPN because the first external dependency happened to need one, then route everything through it out of habit.
The architectural mistake to avoid
Once an IPsec data plane exists, it becomes the path of least resistance for every new integration, because it's already there and it already works. That's how a tunnel built for one legacy appliance ends up carrying your payment provider traffic. Decide per dependency, on what the remote actually requires — not on what you happen to have already built.
StrongSwan sidecar
One pod, one or two tunnels, an Nginx proxy and standard probes. Minimal moving parts and zero new abstractions. Still the right call for a small, stable set of IPsec-only peers.
IPsec operator
Tunnels as custom resources, a controller driving charon over vici, XFRM and VXLAN per connection, pool-based IPAM. Scales IPsec properly — at the cost of a data plane you own.
Kubernetes-native egress
Default-deny policy, an egress gateway with a stable source range, Gateway API routes and identity-based authorisation. The right default for anything that speaks TLS.
This completes the 3-part series on external connectivity from Kubernetes.
Part 1 — IPsec out of Kubernetes with StrongSwan
Part 2 — Managing IPsec tunnels with an operator
Part 3 — External connectivity, the Kubernetes-native way (you're reading it)
