Skip to content Skip to footer

External Connectivity, the Kubernetes-Native Way

External Connectivity, the Kubernetes-Native Way

Parts 1 and 2 built increasingly good answers to the same question: how do you get an IPsec tunnel out of a Kubernetes cluster and keep it alive? This post asks a harder one. whether you should have been building a tunnel at all. For a growing share of external dependencies the answer is no, and the cluster already has better primitives for the job.

The question the operator couldn't answer

The operator from Part 2 is genuinely good engineering. Tunnels are declarative, self-healing, observable, and workloads join them with an annotation. If you have a fleet of IPsec peers, build it — nothing in this post retracts that.

But step back and look at what it costs to run. You own an IKE daemon, per-connection XFRM interfaces, a VXLAN overlay, an IPAM system, a mutating webhook, and an MTU problem. That is a small networking product, maintained by a platform team whose actual job is something else. And it's worth asking honestly what fraction of your external dependencies genuinely require it.

For most teams the answer is uncomfortable. The legacy appliance that started this whole series does need IPsec — it speaks nothing else. But the partner API, the payment provider, the object storage endpoint, the observability vendor, the internal service in a peered VPC: those all speak TLS. Routing them through an IKE data plane you hand-built doesn't make them more secure. It makes them dependent on kernel plumbing that has nothing to do with the guarantee anyone actually wanted, which was this identified workload may talk to that specific endpoint, encrypted, and I can prove it happened.

Cluster workloads routed through heavy IPsec machinery toward external dependencies, most of which only need TLS
The machinery from Parts 1 and 2 serves one column honestly. For everything that already speaks TLS, the tunnel is overhead wearing a security costume.

The pattern: make egress a policy, not a tunnel

The Kubernetes-native answer inverts the question. Instead of asking "how do I build a private path to that system," you ask "which workload identity is allowed to reach which destination, over what encryption, and how do I see it." Those are all things the cluster can already express.

Three primitives carry almost all of it. NetworkPolicy gives you default-deny egress, so nothing leaves a namespace unless you said it could. An egress gateway — a dedicated set of proxy replicas on their own nodes — becomes the single funnel through which permitted traffic exits, giving the outside world one stable source address to allowlist. And workload identity, issued as short-lived certificates, means the policy decision is made against who the caller is rather than which IP it happened to get from the CNI.

That last shift is the substantive one. In Part 1, the remote side saw an Nginx proxy's address. In Part 2 it saw an address from a managed pool, which was better. Here it sees a stable egress IP and the call itself carries a cryptographic identity the mesh verified before it left. The IP allowlist stops being your security model and becomes what it should always have been: a coarse network control layered under a real one.

Egress gateway architecture with default-deny policy, identity-based routing, and a stable egress IP range
Default-deny at the namespace, identity-based permission in the policy layer, and one stable egress range the partner can allowlist. No pod knows a gateway exists.

Why this is different from just "put a proxy in front of it"

Part 1's Nginx was also a proxy, and it was a real hop that erased the client. The difference isn't the proxy — it's that the egress gateway is told what to do by declarative objects rather than a config file baked into an image, and that it makes decisions on verified identity rather than on the source IP of whatever pod happened to connect. Same topology, completely different operational model.

Gateway API, pointed outward

Most platform teams already know the Gateway API from ingress. The useful realisation is that the same object model describes egress: a Gateway with a listener, routes that match destinations, and policies attached to backends. You are not learning a new abstraction — you are pointing an existing one in the other direction, which means the same review process, the same GitOps repo, and the same mental model your team already has.

A permitted external dependency becomes a few objects rather than a tunnel:

# 1 · Default-deny egress for the namespace. Nothing leaves unless allowed. apiVersion: networking.k8s.io/v1 kind: NetworkPolicy metadata: { name: deny-all-egress, namespace: payments } spec: podSelector: {} policyTypes: [Egress] --- # 2 · One route out, matched by hostname, sent through the egress gateway. apiVersion: gateway.networking.k8s.io/v1alpha2 kind: TLSRoute metadata: { name: partner-billing-egress, namespace: payments } spec: parentRefs: [{ name: egress-gateway, namespace: egress }] hostnames: ["billing.partner.example"] rules: - backendRefs: [{ name: partner-billing, port: 443 }] --- # 3 · Verify the far end properly: pin the CA and the SNI, don't just trust TLS. apiVersion: gateway.networking.k8s.io/v1alpha3 kind: BackendTLSPolicy metadata: { name: partner-billing-tls, namespace: payments } spec: targetRefs: [{ kind: Service, name: partner-billing }] validation: hostname: billing.partner.example # SNI + cert name check caCertificateRefs: [{ kind: ConfigMap, name: partner-ca }]

Compare that to Part 2's IPsecTunnel. Both are declarative, both live in Git, both are reconciled by a controller. The difference is what happens underneath: the operator programmed kernel state you own and debug, while this programs a userspace proxy the platform already runs for ingress. The blast radius of a mistake is a rejected connection, not a broken route table on a node.

Gateway API objects for egress reconciled by a controller into Envoy proxy configuration
GatewayClass, Gateway, Route and BackendTLSPolicy are reconciled into proxy configuration over xDS — the same declarative loop as ingress, aimed outward.

Anatomy of a native egress path

It's worth walking a single request end to end, because the interesting work happens in places the application never sees.

1 · Identity — the workload presents a short-lived certificate naming who it is, not where it is
2 · Admission — default-deny holds unless a route and policy explicitly permit this identity to this host
3 · Egress — traffic exits through gateway replicas on dedicated nodes with a stable source range
4 · Verification — the proxy validates the remote certificate against a pinned CA and expected SNI
5 · Telemetry — the call is counted, traced, and logged with the calling identity attached

Step 5 is the one that quietly justifies the whole exercise. With a tunnel, "who called the partner API and when?" is answerable only if the application logged it. Here it falls out of the data path: request counts and latency per identity and destination, TLS handshake failures, certificate expiry, policy denials. You alert on a metric before the partner emails you, and when someone asks which services depend on a vendor you are about to drop, you query it instead of grepping code.

Workload identity issuance, policy decision at the egress proxy, and telemetry flowing to monitoring
Short-lived identities are issued and rotated by the control plane; every outbound call is authorised against them and emits telemetry — including the denials.

The payoff

Adding an external dependency becomes a pull request against a route, reviewed like any other change, with no kernel state, no address pool, and no MTU maths. Removing one is deleting an object. And for the first time the question "what does this cluster talk to on the outside, and who is allowed to?" has an answer that lives in the platform rather than in a wiki someone stopped updating in 2024.

Where the native approach doesn't reach

This is the part the vendor diagrams leave out. Native egress is the right default, not a universal replacement, and pretending otherwise is how platform teams end up rebuilding a tunnel badly at the worst possible moment.

  • The remote may simply not speak TLS. The appliance from Part 1 wants ESP and IKEv2 and will not be talked out of it. No amount of Gateway API changes that — you still need Pattern A or B for that peer.
  • Non-HTTP protocols get awkward. A raw database connection or a legacy binary protocol can be passed through at L4, but you lose most of the routing and observability that made the approach attractive in the first place.
  • The remote may demand a fixed private address. Some partners allowlist an RFC1918 source that only exists inside their VPN. A stable public egress range doesn't satisfy that requirement.
  • You've traded kernel complexity for control-plane complexity. A mesh or Gateway controller is not free. It's a well-supported product rather than bespoke code, which is a real improvement, but it is still a thing to run, upgrade, and debug.

The failure mode to design against: silent expiry

Native egress leans on certificates — the mesh's own short-lived identities, and the CA bundle you pinned for the remote. Both expire. A partner rotating their CA without telling you produces a total, instant outage for exactly one dependency, and the error surfaces as a TLS handshake failure buried in proxy logs. Alert on remote certificate expiry and on handshake-failure rate per destination from day one. This is the native equivalent of Part 2's MTU trap: rare, sharp, and obvious only in hindsight.

Choosing between the three patterns

Three posts, three answers, and the useful conclusion is not that the last one wins. They're a ladder, and the discipline is picking the lowest rung that actually solves your problem.

Comparison of the sidecar, operator, and Kubernetes-native egress patterns with guidance on when to use each
The patterns are a ladder, not a ranking. Most mature platforms end up running Pattern C for the majority of egress and Pattern A or B for the handful of peers that genuinely demand IPsec.
QuestionSidecar (A)Operator (B)Native egress (C)
Remote speaksIPsec onlyIPsec onlyTLS / HTTP
Scales to many peersNo — linear costYesYes
You own kernel stateSome (one pod)Yes — xfrm, VXLAN, IPAMNo
Identity-aware policyNoNo — network-level onlyYes
Per-call observabilityProcess probes onlyTunnel-level metricsPer request and identity
Sharp edge to plan forProxy hides client IPMTU with VXLAN + ESPCertificate expiry
Reach for it when1–2 fixed IPsec peersA fleet of IPsec peersEverything else — the default

The honest end state for most platforms is a mix. Run native egress as the default path for the long tail of TLS dependencies, where it gives you policy, identity and telemetry for far less operational weight. Keep an IPsec capability — a sidecar if it's one or two peers, the operator if it's many — for the appliances that will never change. What you should not do is what the series opened with: build a VPN because the first external dependency happened to need one, then route everything through it out of habit.

The architectural mistake to avoid

Once an IPsec data plane exists, it becomes the path of least resistance for every new integration, because it's already there and it already works. That's how a tunnel built for one legacy appliance ends up carrying your payment provider traffic. Decide per dependency, on what the remote actually requires — not on what you happen to have already built.

Pattern A · Part 1

StrongSwan sidecar

One pod, one or two tunnels, an Nginx proxy and standard probes. Minimal moving parts and zero new abstractions. Still the right call for a small, stable set of IPsec-only peers.

Pattern B · Part 2

IPsec operator

Tunnels as custom resources, a controller driving charon over vici, XFRM and VXLAN per connection, pool-based IPAM. Scales IPsec properly — at the cost of a data plane you own.

Pattern C · this post

Kubernetes-native egress

Default-deny policy, an egress gateway with a stable source range, Gateway API routes and identity-based authorisation. The right default for anything that speaks TLS.


This completes the 3-part series on external connectivity from Kubernetes.
Part 1 — IPsec out of Kubernetes with StrongSwan
Part 2 — Managing IPsec tunnels with an operator
Part 3 — External connectivity, the Kubernetes-native way (you're reading it)

Leave a Comment