A service mesh is one of those tools teams adopt because a conference talk made it look inevitable, then spend a year operating something they didn't strictly need. Istio is also, done right, the cleanest answer to a set of problems every growing platform eventually hits. This first part of the series is about telling those two situations apart — what Istio actually is under the hood, what it solves, and when the honest answer is "not yet".
The problem a mesh exists to solve
Start with a platform of a dozen services and no mesh. Every one of them talks to others over the network, and every team has quietly made its own decisions about how. The orders service in Go loads TLS certificates from a mounted secret and retries three times. Payments in Java has its own TLS setup, different timeouts and a different metrics library. Inventory in Node speaks plaintext because nobody got around to it.
None of that is negligence. It's what happens when networking concerns live inside application code: encryption, retries, timeouts, identity, telemetry — each re-implemented per language, per team, per release cycle. The result is a platform where you cannot answer basic questions with confidence. Is all service-to-service traffic encrypted? Which service is allowed to call payments? Why did latency spike between orders and inventory at 14:02? The answers are scattered across a dozen codebases.
A service mesh moves those concerns out of the application and into the platform. The code makes a plain call; the infrastructure around it handles encryption, identity, policy, retries and telemetry, identically for every workload regardless of language. That's the whole idea. Everything else is implementation.
Istio's architecture: a control plane and a data plane
Istio splits cleanly into two halves, and understanding that split is most of understanding Istio.
The control plane is a single component: istiod. It watches the Kubernetes API for Services, endpoints and Istio's own custom resources, turns that desired state into proxy configuration, and pushes it out over xDS — the same dynamic configuration API Envoy uses everywhere. It also runs a certificate authority that issues every workload a short-lived certificate carrying a SPIFFE identity derived from its Kubernetes service account, and rotates those certificates automatically. That identity is what later lets you write policy about which service may call which, instead of which IP.
The data plane is where traffic actually flows: a fleet of proxies that intercept service-to-service connections, apply the configuration istiod pushed, encrypt with mutual TLS, and emit metrics, traces and access logs. Crucially, istiod is never in the request path. If the control plane goes down, existing traffic keeps flowing with the last known configuration; what stops is new configuration and certificate issuance.
What you get the moment workloads join the mesh
Before writing a single policy: mutual TLS between meshed workloads, a cryptographic identity per service account, and consistent golden-signal metrics — request rate, error rate, latency — for every hop. For many teams that baseline alone justifies the mesh. For others it's exactly the part they already have, which matters for the decision at the end.
Two data plane modes: sidecar and ambient
Where the proxies run is the biggest architectural decision in Istio today, because it now has two answers.
Sidecar mode: a full proxy in every pod
The original model injects an Envoy sidecar into every pod. The application talks to localhost; the sidecar intercepts the connection, applies routing and policy, and speaks mTLS to the sidecar on the other side. Every pod gets full Layer 7 capability — HTTP-aware routing, retries, header-based policy, per-request telemetry — with no shared component between workloads.
The price is paid per pod. Every replica carries an extra container with its own CPU and memory reservation, every rollout of Istio means restarting workloads to pick up the new proxy, and every pod's configuration grows with the size of the mesh unless you scope it. At a few dozen pods none of that matters. At several thousand it becomes a real line in the capacity plan.
Ambient mode: split L4 from L7
Ambient mode, generally available since Istio 1.24, removes the sidecar entirely and splits the proxy's job in two. A lightweight, Rust-based ztunnel runs once per node and handles the Layer 4 essentials for every pod on that node: mTLS, workload identity, L4 authorization and TCP telemetry. Traffic between ztunnels travels over HBONE, an mTLS tunnel built on HTTP CONNECT.
Layer 7 features are opt-in. Where a namespace or service genuinely needs HTTP-aware routing or policy, you deploy a waypoint proxy — a regular Envoy deployment that traffic for that destination passes through. Everything else stays on the cheap L4 path. Pods join the mesh by label, with no injection and no restart, and upgrading the data plane no longer means rolling every workload.
| Aspect | Sidecar mode | Ambient mode |
|---|---|---|
| Proxy placement | One Envoy per pod | ztunnel per node, waypoints where needed |
| L7 features | Everywhere, always | Only behind a waypoint |
| Resource overhead | Scales with pod count | Scales with nodes and chosen waypoints |
| Joining the mesh | Injection + pod restart | Namespace label, no restart |
| Data plane upgrades | Restart workloads | Roll ztunnel and waypoints |
| Maturity | Longest production history | GA, newer — check feature status for edge cases |
Ambient isn't "sidecar but cheaper"
The L4/L7 split changes how you reason about the mesh. In sidecar mode, the client's proxy enforces L7 policy; in ambient, the waypoint in front of the destination does. Policies written for one model don't always mean the same thing in the other, and debugging moves from "look at the pod's sidecar" to "which ztunnel and which waypoint did this hop cross?" Pick a mode deliberately — both can coexist in one mesh, but you want a default, not an accident.
When you don't need a mesh
This is the section most Istio content skips. A mesh is a platform product your team will run, upgrade and debug for years. It earns that cost when the problems above are real for you — and plenty of healthy platforms don't have them yet.
- A handful of services owned by one team. Consistent libraries and conventions are cheaper than a control plane when the same people write every service.
- Encryption is a compliance checkbox, not an identity requirement. Application TLS with certificates from cert-manager, plus default-deny NetworkPolicy, covers "encrypted in transit and segmented" without a mesh.
- You need traffic splitting at the edge only. Canaries for externally facing services can be handled by your ingress or Gateway API implementation.
- Nobody owns platform operations. An unowned mesh is worse than no mesh — expired certificates, stale proxies and mystery 503s land on whoever is on call.
The signals that you do need one are just as concrete: a security requirement that every service-to-service call is authenticated by workload identity, polyglot teams you can't standardise on one library, a need for consistent retries and timeouts across services you don't control, or incidents where nobody can see what happened between two services.
The practical default in 2026
For a new adoption, ambient mode's L4 layer is the lowest-cost way to get mTLS and workload identity across a whole cluster, with waypoints added only where Layer 7 earns its keep. Sidecar mode remains the right call when you need full L7 on most workloads, depend on a feature that's sidecar-only in your Istio version, or already run sidecars successfully. Either way, the decision should come from your requirements — not from which mode a tutorial happened to use.
istiod
Watches the cluster, pushes proxy configuration over xDS, and issues short-lived workload identities. Never in the request path.
Envoy in every pod
Full Layer 7 everywhere and the longest production history — at a per-pod cost in resources and restarts.
ztunnel + waypoints
mTLS and identity per node, Layer 7 only where you place a waypoint. Lower overhead, a different mental model.
This is Part 1 of a series on running Istio in production.
Part 1 — Istio Service Mesh Architecture: What a Mesh Actually Solves (you're reading it)
Part 2 — Rolling Istio into a live cluster, namespace by namespace (coming next)
Part 3 — Istio traffic management: canaries, retries and circuit breaking
Part 4 — Zero-trust with Istio: mTLS, workload identity and AuthorizationPolicy
