Rate Limiting for Workloads Within Istio
I recently rolled out a rate-limiting platform across a shared Kubernetes cluster. In theory, it sounds simple: follow the upstream Istio docs and deploy Envoy’s reference ratelimit service. But as always, the devil is in the operational details.
By 2026, most mid-to-large engineering orgs run an Internal Developer Platform (IDP). Plain Helm is great, but it falls short when enforcing standardization, governance, and self-service deployment workflows across hundreds of business services. Rate limiting couldn't be a standalone bespoke component; it had to be a natural extension of our existing IDP.
That meant extending our onboarding contracts, exposing clean abstractions to product teams, and reconciling user intent with the underlying mesh state. The architecture naturally broke down into three tiers: the IDP, a custom Control Plane, and the Rate Limit Service (RLS).
flowchart TD
IDP-CLI -. Custom CRD .-> cp[Rate Limit<br>Control Plane]
IDP-CLI -. Standard Deployment .-> app1[App Pod]
cp -. Push Config .-> rls[Rate Limit Service]
User --> app1
app1 -- Check Limit --> rlsWhat I Anticipated (and Prepared For)
Asynchronous compilation and health status
A custom Kubernetes Control Plane (operator) can perform validation, but applying a CRD successfully to the Kubernetes API doesn't mean the policy is live.
Envoy's rate limit service needs to compile the configuration snapshot and load it into memory. If that policy fails to compile or the RLS rejects the snapshot, the CRD status must reflect it. Putting an error message in status.conditions is standard practice, but nobody reads status fields during routine deployments.
Because reconciliation happens asynchronously—often a few seconds after the manifest applies—CI/CD pipelines would report a successful deploy even if the policy failed downstream. Fortunately, Argo CD supports custom Lua health checks. Writing a custom Lua script allowed Argo to evaluate the subresource status directly, keeping the application in a Progressing or Degraded state until the control plane explicitly marked the compiled snapshot as healthy.
What I Wish I’d Addressed Earlier
1. Ubiquitous language vs. platform jargon
As a recent joiner, I underestimated the stickiness of the company's internal vocabulary. Many organizations migrating legacy architectures to Kubernetes carry legacy domain terms with them.
Whether an internal term makes technical sense or not, if every team uses it, fighting it is a losing battle. I initially designed the CRD specs and IDP definitions around simple terminology. Once the control plane and workloads were running, the mismatch became obvious, forcing translation layers in the CLI and late retrofits. Adopting local conventions upfront would have eliminated refactor which followed after.
2. Ephemeral environments and volatile resources
A core early assumption was that our Control Plane operator didn't need sub-second reconciliation. Whether a policy reflected 0.5 seconds or 5 seconds after an Argo sync felt negligible.
Then came the edge cases. The IDP team actively load-tested and hammered the platform to diagnose some race conditions. In fast-paced dynamic environments, resources don't stay still:
- An app deployment might be deleted 2 seconds after its rate-limit CRD is applied.
- Entire test namespaces get spun up, modified, and torn down in rapid succession.
The operator originally assumed target workloads would exist throughout the reconciliation loop. When a namespace vanished mid-reconciliation, unhandled NotFound errors triggered alerts in development. Operators must be resilient to aggressive resource churn or alerts will fire.
3. Latency budgets, cold starts, and node evictions
Pod evictions, Karpenter node consolidations, and spot interruptions are standard Kubernetes behavior. What made rate limiting tricky was the strict latency budget.
Because the rate limit check sits directly in the request path, the client-side timeout must be tiny—often just a few milliseconds. During cold starts response latency briefly spikes.
Depending on your configuration:
- Fail-open (
failure_mode_deny: false): Brief bursts of unthrottled traffic pass through during hiccups. - Fail-close (
failure_mode_deny: true): Legitimate requests fail with500or429errors.
Even a fraction of a percent (0.0x%) of traffic dropping or bypassing limits won't turn your core dashboard red, but it generates an annoying amber flicker that erodes confidence.