Skip to content Skip to footer

Istio Service Mesh Architecture – Part1

Istio Service Mesh Architecture - Part1

A service mesh is one of those tools teams adopt because a conference talk made it look inevitable, then spend a year operating something they didn't strictly need. Istio is also, done right, the cleanest answer to a set of problems every growing platform eventually hits. This first part of the series is about telling those two situations apart — what Istio actually is under the hood, what it solves, and when the honest answer is "not yet".

The problem a mesh exists to solve

Start with a platform of a dozen services and no mesh. Every one of them talks to others over the network, and every team has quietly made its own decisions about how. The orders service in Go loads TLS certificates from a mounted secret and retries three times. Payments in Java has its own TLS setup, different timeouts and a different metrics library. Inventory in Node speaks plaintext because nobody got around to it.

None of that is negligence. It's what happens when networking concerns live inside application code: encryption, retries, timeouts, identity, telemetry — each re-implemented per language, per team, per release cycle. The result is a platform where you cannot answer basic questions with confidence. Is all service-to-service traffic encrypted? Which service is allowed to call payments? Why did latency spike between orders and inventory at 14:02? The answers are scattered across a dozen codebases.

Three services in different languages, each carrying its own TLS, retry and metrics logic, with inconsistent encryption between them
Without a mesh, every service owns its own networking stack — and the guarantees across the platform are only as strong as the weakest one.

A service mesh moves those concerns out of the application and into the platform. The code makes a plain call; the infrastructure around it handles encryption, identity, policy, retries and telemetry, identically for every workload regardless of language. That's the whole idea. Everything else is implementation.

Istio's architecture: a control plane and a data plane

Istio splits cleanly into two halves, and understanding that split is most of understanding Istio.

The control plane is a single component: istiod. It watches the Kubernetes API for Services, endpoints and Istio's own custom resources, turns that desired state into proxy configuration, and pushes it out over xDS — the same dynamic configuration API Envoy uses everywhere. It also runs a certificate authority that issues every workload a short-lived certificate carrying a SPIFFE identity derived from its Kubernetes service account, and rotates those certificates automatically. That identity is what later lets you write policy about which service may call which, instead of which IP.

The data plane is where traffic actually flows: a fleet of proxies that intercept service-to-service connections, apply the configuration istiod pushed, encrypt with mutual TLS, and emit metrics, traces and access logs. Crucially, istiod is never in the request path. If the control plane goes down, existing traffic keeps flowing with the last known configuration; what stops is new configuration and certificate issuance.

Kubernetes API watched by istiod, which pushes xDS configuration to data-plane proxies and signs short-lived workload certificates, while proxies emit telemetry
istiod turns cluster state into proxy configuration and workload identities. The proxies do the work; the control plane never touches a request.

What you get the moment workloads join the mesh

Before writing a single policy: mutual TLS between meshed workloads, a cryptographic identity per service account, and consistent golden-signal metrics — request rate, error rate, latency — for every hop. For many teams that baseline alone justifies the mesh. For others it's exactly the part they already have, which matters for the decision at the end.

Two data plane modes: sidecar and ambient

Where the proxies run is the biggest architectural decision in Istio today, because it now has two answers.

Sidecar mode: a full proxy in every pod

The original model injects an Envoy sidecar into every pod. The application talks to localhost; the sidecar intercepts the connection, applies routing and policy, and speaks mTLS to the sidecar on the other side. Every pod gets full Layer 7 capability — HTTP-aware routing, retries, header-based policy, per-request telemetry — with no shared component between workloads.

The price is paid per pod. Every replica carries an extra container with its own CPU and memory reservation, every rollout of Istio means restarting workloads to pick up the new proxy, and every pod's configuration grows with the size of the mesh unless you scope it. At a few dozen pods none of that matters. At several thousand it becomes a real line in the capacity plan.

Two pods, each with an application container and an Envoy sidecar; the sidecars exchange mTLS traffic and each receives xDS configuration from istiod
Sidecar mode: maximum capability per pod, paid for with one proxy per replica and workload restarts on every data-plane upgrade.

Ambient mode: split L4 from L7

Ambient mode, generally available since Istio 1.24, removes the sidecar entirely and splits the proxy's job in two. A lightweight, Rust-based ztunnel runs once per node and handles the Layer 4 essentials for every pod on that node: mTLS, workload identity, L4 authorization and TCP telemetry. Traffic between ztunnels travels over HBONE, an mTLS tunnel built on HTTP CONNECT.

Layer 7 features are opt-in. Where a namespace or service genuinely needs HTTP-aware routing or policy, you deploy a waypoint proxy — a regular Envoy deployment that traffic for that destination passes through. Everything else stays on the cheap L4 path. Pods join the mesh by label, with no injection and no restart, and upgrading the data plane no longer means rolling every workload.

Pods without sidecars on two nodes; each node runs a ztunnel handling L4 mTLS over HBONE, with an optional namespace waypoint proxy for L7 policy and routing
Ambient mode: every pod gets mTLS and identity from its node's ztunnel; only the traffic that needs Layer 7 pays for a waypoint.
AspectSidecar modeAmbient mode
Proxy placementOne Envoy per podztunnel per node, waypoints where needed
L7 featuresEverywhere, alwaysOnly behind a waypoint
Resource overheadScales with pod countScales with nodes and chosen waypoints
Joining the meshInjection + pod restartNamespace label, no restart
Data plane upgradesRestart workloadsRoll ztunnel and waypoints
MaturityLongest production historyGA, newer — check feature status for edge cases

Ambient isn't "sidecar but cheaper"

The L4/L7 split changes how you reason about the mesh. In sidecar mode, the client's proxy enforces L7 policy; in ambient, the waypoint in front of the destination does. Policies written for one model don't always mean the same thing in the other, and debugging moves from "look at the pod's sidecar" to "which ztunnel and which waypoint did this hop cross?" Pick a mode deliberately — both can coexist in one mesh, but you want a default, not an accident.

When you don't need a mesh

This is the section most Istio content skips. A mesh is a platform product your team will run, upgrade and debug for years. It earns that cost when the problems above are real for you — and plenty of healthy platforms don't have them yet.

  • A handful of services owned by one team. Consistent libraries and conventions are cheaper than a control plane when the same people write every service.
  • Encryption is a compliance checkbox, not an identity requirement. Application TLS with certificates from cert-manager, plus default-deny NetworkPolicy, covers "encrypted in transit and segmented" without a mesh.
  • You need traffic splitting at the edge only. Canaries for externally facing services can be handled by your ingress or Gateway API implementation.
  • Nobody owns platform operations. An unowned mesh is worse than no mesh — expired certificates, stale proxies and mystery 503s land on whoever is on call.

The signals that you do need one are just as concrete: a security requirement that every service-to-service call is authenticated by workload identity, polyglot teams you can't standardise on one library, a need for consistent retries and timeouts across services you don't control, or incidents where nobody can see what happened between two services.

Decision flow: no identity-based encryption needed leads to no mesh; identity but no L7 leads to ambient L4 only; selective L7 leads to ambient with waypoints; broad L7 leads to sidecar mode
Pick the lowest rung that solves your actual problem. Each step up adds capability — and something new to operate.

The practical default in 2026

For a new adoption, ambient mode's L4 layer is the lowest-cost way to get mTLS and workload identity across a whole cluster, with waypoints added only where Layer 7 earns its keep. Sidecar mode remains the right call when you need full L7 on most workloads, depend on a feature that's sidecar-only in your Istio version, or already run sidecars successfully. Either way, the decision should come from your requirements — not from which mode a tutorial happened to use.

Control plane

istiod

Watches the cluster, pushes proxy configuration over xDS, and issues short-lived workload identities. Never in the request path.

Data plane · sidecar

Envoy in every pod

Full Layer 7 everywhere and the longest production history — at a per-pod cost in resources and restarts.

Data plane · ambient

ztunnel + waypoints

mTLS and identity per node, Layer 7 only where you place a waypoint. Lower overhead, a different mental model.


This is Part 1 of a series on running Istio in production.
Part 1 — Istio Service Mesh Architecture: What a Mesh Actually Solves (you're reading it)
Part 2 — Rolling Istio into a live cluster, namespace by namespace (coming next)
Part 3 — Istio traffic management: canaries, retries and circuit breaking
Part 4 — Zero-trust with Istio: mTLS, workload identity and AuthorizationPolicy

Leave a Comment