Chaos Engineering on Kubernetes: Chaos Mesh, LitmusChaos and Steady-State Hypotheses
Most Kubernetes outages are not caused by a component failing. They are caused by a component failing in a way nobody believed it could, while a retry policy, a readiness probe, or a pod disruption budget quietly did the opposite of what its author intended. The cluster looked redundant on a diagram, and the first real failure proved otherwise. That is the gap chaos engineering Kubernetes practice exists to close: you inject the failure on purpose, at a time you choose, with a way to stop, and you learn whether the system’s resilience is real or imagined.
This matters more in 2026 than it did when the discipline was young. Clusters now host stateful databases, message brokers, edge gateways and agentic control loops, and the dependency graph behind a single request often crosses a dozen services. You will leave this post with a working mental model of the experiment lifecycle, a fair comparison of the two CNCF chaos projects, runnable YAML for a pod fault and a network fault, and a plan for blast-radius control, CI integration and your first game day.
What this covers: the steady-state hypothesis method, the Chaos Mesh and LitmusChaos architectures, a safe experiment walk-through, abort conditions tied to SLOs, RBAC and scope limits, managed alternatives, and the failure modes of chaos tooling itself.
Context and Background
Chaos engineering was formalised at Netflix as a discipline of experimenting on a distributed system to build confidence in its ability to withstand turbulent conditions. The community-maintained Principles of Chaos Engineering distil the method into a hypothesis-driven loop: define steady state as a measurable output, hypothesise that it persists in both a control group and an experimental group, introduce variables that mimic real-world disruptions, and then try to disprove the hypothesis. The same document lists five advanced principles: build a hypothesis around steady-state behaviour, vary real-world events, run experiments in production, automate experiments to run continuously, and minimise blast radius.
Notice what that list does not say. It does not say “break things randomly.” The method is closer to a controlled experiment in science than to a stress test. You choose a fault, predict what the system will do, observe, and treat any difference between prediction and observation as a finding. A pod kill that changes nothing is a successful experiment because it confirms a belief you can now stop worrying about.
Kubernetes changes the practice in three ways. First, the platform already performs a degree of self-healing, so the interesting failures are the ones that defeat or confuse that machinery: a pod that is Running but not serving, a node that is partitioned but not NotReady, a DNS lookup that takes four seconds. Second, everything is declarative, which makes faults themselves expressible as custom resources that live in Git, get reviewed, and run from pipelines. Third, multi-tenancy means a badly scoped experiment can hit a neighbour’s namespace, so scope control is part of the tool, not an afterthought.
Two open-source projects dominate this space, and both are CNCF projects. Chaos Mesh, originally built by PingCAP, was accepted as a sandbox project and moved to incubating in early 2022. LitmusChaos entered the CNCF sandbox in 2020 and was promoted to incubating in January 2022. Both adopt a Kubernetes-native approach in which chaos intent is declared as custom resources. Beyond them sit commercial platforms such as Gremlin, and cloud-managed services such as AWS Fault Injection Service and Azure Chaos Studio, which, as we will see, can orchestrate the open-source tools rather than replace them.
If you are building on a platform that also hosts industrial or edge workloads, the network layer deserves the same scrutiny as the workloads. Our Cilium vs Calico ADR for industrial Kubernetes networking covers the CNI trade-offs that determine which network faults you can even inject cleanly.
The Experiment Lifecycle: A Reference Architecture for Chaos on Kubernetes
Every chaos experiment on Kubernetes, regardless of tool, follows the same six-stage loop: define a steady-state metric, declare a hypothesis with a numeric tolerance, scope the blast radius, inject the fault through a custom resource, watch the metric against abort conditions, and then record findings and fix. The tooling automates injection and abort; the thinking at stages one to three is the real work.

Figure 1: The chaos engineering Kubernetes experiment lifecycle, from steady-state definition through injection, observation and remediation.
The diagram shows the loop as a closed circuit with an explicit abort edge. The observation stage reads SLO metrics from your monitoring stack and compares them to the hypothesis. If the metric breaches the abort threshold, the controller deletes the fault resource, which triggers recovery. If the experiment runs to completion within tolerance, the finding is “hypothesis held,” and the loop restarts with a more severe fault. The loop closes only when a finding becomes a ticket, a fix, and a regression experiment.
Stage one: define steady state as an externally visible output
A steady-state hypothesis must be about what a user or caller experiences, not about internals. “CPU is below 70 percent” is not steady state; it is a resource measurement. “The checkout API returns a success response within 300 milliseconds for 99 percent of requests over a rolling five-minute window” is steady state. The distinction matters because internals can change wildly while the output remains healthy, and the reverse. The Principles of Chaos Engineering make the same point: measure throughput, error rates and latency percentiles rather than internal mechanics.
In practice, pick two or three signals that already exist as service level indicators. A request success ratio, a latency percentile and, for pipelines, a consumer lag or freshness figure cover most services. If you do not have these signals, you do not yet have a chaos engineering problem. You have an observability problem, and the first experiment should be to kill a pod and see whether anything in your dashboards notices.
Stage two: write the hypothesis with a tolerance and a duration
A good hypothesis has four parts: the fault, the scope, the expected steady-state behaviour and the tolerance. For example: “When one of three replicas of the cart service is killed, the success ratio of the cart API stays at or above 99.5 percent and p99 latency stays below 400 milliseconds, with recovery to full replica count within 60 seconds.” Every clause is falsifiable. If the success ratio dips to 97 percent for twelve seconds, the hypothesis is disproved, and you have learned something specific, such as that clients hold connections to a terminated endpoint until a timeout expires.
Writing the tolerance forces a conversation that is usually skipped: what degradation is acceptable? Teams frequently discover that they have no agreed answer, and that the real SLO lives in the head of the person on call. That discovery is itself a deliverable.
Stage three: constrain the blast radius before you inject anything
Blast radius is the set of users, services and data that a fault can affect. You control it along four axes: which resources the fault selects (namespace, labels, a count or percentage of pods), how long the fault lasts, how much of the traffic it touches, and which environment it runs in. A mature programme starts with a single replica in a staging namespace for thirty seconds, and expands one axis at a time. Expanding two at once makes a failed experiment hard to interpret.
The Kubernetes-native tools express these axes directly. Chaos Mesh uses a selector block with namespaces and label selectors, a mode that picks one, all, a fixed number, or a percentage of matching pods, and a duration. LitmusChaos expresses the same ideas through experiment environment variables such as target labels, chaos duration and pods-affected percentage. Section four looks at both in detail.
What the controller actually does
Under the hood, both tools run a controller that watches fault custom resources and a privileged per-node component that performs the injection. For network and I/O faults, that component enters the target container’s namespaces or manipulates the traffic-control and filesystem layers of the node. This is why chaos tooling needs elevated permissions, and why its installation is a security decision. We return to this in the gotchas section. For now, hold the model: a declarative resource says what to break, a controller reconciles it, a node agent breaks it, and the same controller undoes it when the duration expires or the resource is deleted. The delete-to-recover property is what makes abort conditions possible, because aborting an experiment is simply removing an object.
Chaos Mesh vs LitmusChaos: Architectures and Fault Catalogues
Chaos Mesh and LitmusChaos solve the same problem with different centres of gravity. Chaos Mesh is a fault-injection engine with a dashboard and workflow layer, where each fault type is its own custom resource. LitmusChaos is a chaos platform with a control plane (ChaosCenter), a library of experiments (ChaosHub), and resilience probes that judge each run. Choose Chaos Mesh for fine-grained fault primitives, and Litmus when you want experiment orchestration and pass/fail verdicts built in.

Figure 2: Component view of Chaos Mesh and LitmusChaos on a Kubernetes cluster, showing where faults are declared, reconciled and injected.
The figure places both tools on one cluster. In Chaos Mesh, a controller manager reconciles fault custom resources and instructs a chaos daemon that runs on every node, while an optional dashboard and a workflow engine sit above. In LitmusChaos, a ChaosCenter control plane stores experiments and results, and a lightweight infrastructure agent in each target cluster pulls experiment definitions, launches experiment pods that run the fault, and reports probe outcomes back.
Chaos Mesh in detail
At the time of writing, the Chaos Mesh supported releases page lists 2.8.4 as the current line, with 2.7.3 and 2.6.7 also supported. The project notes that it replaced an earlier plan of a new release every three months with a roughly six-month cadence, because it lacks enough maintainers for the faster rhythm. That is a useful signal for planning: pin the version, test upgrades in a non-production cluster, and do not expect rapid feature turnaround. The same page says Kubernetes 1.22 support begins with 2.0.4, and that 2.6 is expected to work on 1.26 through 1.28 although the project does not run end-to-end tests on those versions. Check the page for the version matrix that applies to your cluster before installing.
The fault catalogue is organised as one custom resource kind per domain:
- PodChaos offers
pod-failure,pod-killandcontainer-kill. Pod failure makes a pod unavailable for a duration, while pod kill deletes it and relies on a ReplicaSet or similar to recreate it. - NetworkChaos offers
partition,netem(latency, loss, reordering, corruption),bandwidthanddelaystyle emulation. Partitions support a direction ofto,fromorbothand a separatetargetselector for the other side. - StressChaos exerts CPU and memory pressure on containers.
- IOChaos injects latency, errors and attribute faults into filesystem calls on a volume.
- TimeChaos skews the clock seen by target processes, which is invaluable for exposing certificate-expiry and token-lifetime bugs.
- DNSChaos, HTTPChaos, KernelChaos, JVMChaos and PhysicalMachineChaos cover higher-level and non-Kubernetes targets. The exact set and maturity of each varies by release, so confirm against the docs for your pinned version.
Chaos Mesh also ships a Workflow resource that chains faults serially or in parallel with deadline-based steps, and a dashboard that can create, schedule and visualise experiments. Its documentation warns that for network emulation the controller manager must be able to reach the chaos daemon during injection; otherwise injected faults cannot be recovered. That is a real operational hazard and a good reason to practise recovery as part of the first game day.
LitmusChaos in detail
LitmusChaos describes itself as a cloud-native chaos engineering framework that takes a declarative, Kubernetes-native approach. Its building blocks are chaos experiments, which you pick from a hub or author yourself, and chaos scenarios (called workflows or experiments in the ChaosCenter UI depending on version) that compose several faults, define expected outcomes and let you observe the result. The default fault library, ChaosHub, spans Kubernetes, Linux, AWS, GCP and more.
The distinctive feature is the resilience probe. A probe is a check evaluated before, during or after fault injection: an HTTP probe against an endpoint, a command probe run in a pod, a Kubernetes probe that inspects a resource, or a Prometheus probe that evaluates a query. The experiment passes or fails based on probe outcomes, and a weighted score summarises resilience across a scenario. This maps almost one-to-one onto the steady-state hypothesis: the probe is the steady-state check, and its failure aborts or fails the run.
The CNCF’s first-half 2026 update on LitmusChaos reports monthly releases from 3.25.0 in January to 3.30.0 in June 2026. Notable changes include job targeting in chaos experiments (3.27.0), removal of the 1,024-character limit on command probes (3.27.0), Prometheus metrics support and the ability to stop experiments when infrastructure disconnects (3.29.0), and RBAC restriction on project membership queries (3.30.0). The same article describes a Model Context Protocol integration introduced at KubeCon India, letting engineers interact with chaos tooling in natural language, and states the contributor base grew from 241 to 279. Treat the MCP integration as young; I have not verified its production readiness.
A decision matrix for the two projects
| Criterion | Chaos Mesh | LitmusChaos |
|---|---|---|
| Core model | One CRD per fault domain | Experiments in ChaosHub, run by an infrastructure agent |
| Verdict logic | Bring your own (alerts, scripts, workflow) | Built-in probes and resilience score |
| Control plane | Optional dashboard, no central multi-cluster store | ChaosCenter with projects, RBAC and GitOps sync |
| Strengths | Fine-grained network, I/O and time faults | Orchestration, probes, multi-cluster, non-Kubernetes targets |
| Release cadence | Roughly six-monthly | Monthly in H1 2026 |
| Best fit | Platform teams that script their own checks | Teams wanting a ready pass/fail framework |
Neither is a safe default for everyone. A small platform team that already runs Prometheus alerts and Argo Workflows may get more from Chaos Mesh’s primitives. A product organisation wanting self-service experiments for many teams will likely value ChaosCenter’s projects and probes.
Walk-through: A Safe First Experiment with Chaos Mesh
The safest first experiment is a single pod kill against one replica of a stateless service in a staging namespace, with a Prometheus-based abort condition. It costs almost nothing, tests the most common assumption (that replicas are interchangeable), and exercises the whole loop. This walk-through installs Chaos Mesh with namespace filtering on, runs a PodChaos kill, then a NetworkChaos delay.
Install with namespace filtering enabled
Chaos Mesh’s namespace filter is off by default. According to the project documentation on enabled namespaces, you enable it with a Helm flag, then annotate only the namespaces where chaos is allowed. Every other namespace is protected.
helm repo add chaos-mesh https://charts.chaos-mesh.org
helm repo update
# Install into its own namespace with the namespace filter on.
# Add --set chaosDaemon.runtime=containerd (and the matching socketPath)
# if your cluster does not use the chart default; check your CRI.
helm install chaos-mesh chaos-mesh/chaos-mesh \
--namespace chaos-testing --create-namespace \
--set controllerManager.enableFilterNamespace=true
# Allow chaos ONLY in the staging namespace
kubectl annotate ns shop-staging chaos-mesh.org/inject=enabled
# Verify
kubectl get pods -n chaos-testing
kubectl get ns shop-staging -o jsonpath='{.metadata.annotations}'
Pin the chart version in real use rather than taking the latest, and record it in Git. The filter is the single most valuable guard rail in the whole installation: it means a typo in a selector cannot reach production namespaces.
Experiment one: kill a pod
This PodChaos resource kills one pod matching the label of the cart service. The pod must be managed by a ReplicaSet or similar, or it will not come back.
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
name: cart-pod-kill
namespace: shop-staging
labels:
experiment: cart-resilience
spec:
action: pod-kill
mode: one # pick one matching pod at random
selector:
namespaces:
- shop-staging
labelSelectors:
app: cart
gracePeriod: 0 # 0 = immediate kill, no graceful shutdown
Apply it with kubectl apply -f cart-pod-kill.yaml and watch the success ratio. With gracePeriod: 0 the pod is terminated without running its preStop hook, which simulates a node crash better than a graceful rollout does. If you want to test the graceful path, set a non-zero grace period and compare the two runs. They often differ, because connection draining bugs only appear in one of them.
Experiment two: add latency between two services
Latency is the more revealing fault, because it triggers timeout, retry and circuit-breaker behaviour. The following NetworkChaos delays traffic from the cart pods to their downstream dependency for two minutes.
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: cart-to-inventory-delay
namespace: shop-staging
spec:
action: delay
mode: all
selector:
namespaces:
- shop-staging
labelSelectors:
app: cart
direction: to
target:
mode: all
selector:
namespaces:
- shop-staging
labelSelectors:
app: inventory
delay:
latency: "250ms"
jitter: "50ms"
correlation: "100"
duration: "2m"
Choose the latency relative to your timeouts, not in the abstract. If the cart service times out inventory calls at 200 milliseconds, a 250 millisecond delay turns every call into a timeout, and the interesting question becomes what the fallback does. If the timeout is two seconds, 250 milliseconds is merely a slowdown, which tests queueing and thread-pool saturation instead. Run both. Note that Net Emulation depends on the NET_SCH_NETEM kernel module; most distributions include it, but some, such as CentOS, need the kernel-modules-extra package.
For a partition rather than a delay, use action: partition with direction: both. A partition is a harsher fault than latency for most RPC stacks, because it produces hangs until a TCP or application timeout rather than fast failures.
Tying abort conditions to SLOs
Neither fault above is safe without an automatic way to stop. An abort condition is a rule, evaluated continuously, that ends the experiment when steady state is breached beyond tolerance. The mechanism is simple because of the property noted earlier: removing the fault object triggers recovery.

Figure 3: An SLO-driven abort. A watcher queries Prometheus, and on a breach it deletes the fault resource so the controller restores normal behaviour.
In the sequence, the engineer applies the fault, the controller instructs the node daemon, and an independent watcher polls the steady-state query. When the burn rate or error ratio crosses the threshold, the watcher deletes the custom resource and the controller reverses the injection. The watcher must be independent of the thing under test. If the abort logic lives in the same service or the same network path you are degrading, the experiment can blind its own safety net.
A minimal watcher can be a shell loop in a CI job or a small Kubernetes Job. This one aborts when the five-minute error ratio exceeds two percent. The PromQL is illustrative; adapt metric and label names to your own instrumentation.
#!/usr/bin/env bash
# Abort the experiment if the cart error ratio breaches tolerance.
PROM=http://prometheus.monitoring:9090
Q='sum(rate(http_requests_total{app="cart",code=~"5.."}[1m]))
/ sum(rate(http_requests_total{app="cart"}[1m]))'
for i in $(seq 1 24); do # 24 x 5s = 2 minutes
ratio=$(curl -s --data-urlencode "query=$Q" $PROM/api/v1/query \
| jq -r '.data.result[0].value[1] // "0"')
if awk "BEGIN{exit !($ratio > 0.02)}"; then
echo "SLO breach ($ratio) - aborting"
kubectl delete networkchaos cart-to-inventory-delay -n shop-staging
exit 1
fi
sleep 5
done
echo "Hypothesis held"
Chaos Mesh also supports a Schedule resource for recurring experiments and a Workflow resource with deadlines, but it does not itself judge your SLO. That is the gap LitmusChaos probes fill, and it is why many teams combine a fault engine with an external judge such as Prometheus alert rules or an Argo Rollouts-style analysis step.
The same experiment in LitmusChaos
In LitmusChaos, the equivalent pod-kill is the pod-delete experiment, defined as a ChaosEngine that points at an application by label and carries environment variables for duration, interval and the percentage of pods affected, plus probes that act as the steady-state check. Rather than reproduce a ChaosEngine from memory and risk a version-specific error, I recommend generating one from ChaosCenter for your installed version and then committing the exported manifest to Git. The structural point is what matters: fault, target selection and probe live together in one declared object, and a failing probe fails the run. Probes carry a mode (start of test, end of test, continuous, or edge), which lets you check steady state before injection as well as during it. That pre-check is underrated. An experiment run against an already-degraded system proves nothing.
Blast-Radius Control, RBAC and Safe Defaults
Blast-radius control is the set of technical limits that keep an experiment from exceeding its intended scope. In Kubernetes it has five layers: namespace allow-listing, label selectors with a bounded mode, short durations, RBAC on who may create fault resources, and admission policy that rejects unsafe specs. Apply all five; any single one fails eventually.
Namespace allow-listing and selectors
The Chaos Mesh namespace filter described above is your first line. Use it. In Litmus, restrict the application namespace and service account of each experiment to the target namespace. Within a namespace, prefer a mode that selects one pod or a bounded count over all. Chaos Mesh’s mode: all against a label that unexpectedly matches a shared sidecar or a monitoring pod is a classic way to take down observability during an experiment, which defeats the abort logic.
RBAC for chaos operators
Creating a fault resource is a powerful act, so treat the permission like production write access. The Chaos Mesh dashboard supports Manager and Viewer roles at either cluster or namespace scope. A namespace-scoped Manager gets create, delete, patch and update on the chaos-mesh.org API group in that namespace only, plus read access to pods and namespaces. That is the right default for a product team that owns its namespace. Reserve cluster-scoped Manager for the platform team. Note also that Kubernetes 1.24 and later no longer create long-lived token secrets automatically, so dashboard login tokens require a manually created secret, and you should rotate them.
Separately, think about who can reach the dashboard. A chaos dashboard exposed without authentication is a remote fault-injection service. Keep it on a private ingress or reach it only through port-forwarding, and put it behind your identity provider.
Admission policy and guard rails you can enforce
Use an admission controller, such as Kyverno or OPA Gatekeeper, to codify rules about fault resources: duration must be present and below a ceiling, mode: all is rejected outside designated namespaces, selectors must include a namespace, and production namespaces are denied outright unless an approval label is present. This turns a wiki page of guidelines into an enforced contract. Also ensure the target workloads carry PodDisruptionBudgets, because voluntary-disruption semantics and chaos faults interact: a pod-kill fault deletes pods directly and does not honour PDBs, which is exactly why it reveals whether your replicas can absorb a sudden loss.
Protect the control path
Two components must stay healthy for chaos to be safe: the controller that reverses faults and the node agent. If a network partition severs the controller from the daemon on a node, as the Chaos Mesh documentation warns, recovery can fail. Run the chaos controller in a namespace and on nodes that your own experiments never target, and never select the chaos namespace in a fault. The Chaos Mesh PodChaos documentation also cautions against running the controller manager on the target pod.
Integrating Chaos into CI/CD and Game Days
There are two operating modes. Automated experiments run continuously or per release, and game days are scheduled, human-run exercises where people practise the response. A healthy programme uses both: automation catches regressions in known-good behaviour, and game days test the unknowns, the runbooks, and the people.

Figure 4: How automated chaos stages and human game days fit into a delivery pipeline and feed findings back into the backlog.
The figure shows a pipeline in which each build deploys to an ephemeral or staging namespace, runs a short, bounded chaos suite with SLO-based abort, and gates promotion on the verdict. Longer and harsher experiments run on a schedule against a long-lived staging environment, and game days feed findings into the backlog and create new automated regression checks.
Chaos in the pipeline
For per-release checks, keep the suite small: one pod-kill, one latency fault and one dependency-unavailable partition per critical service, each under two minutes. The aim is a regression signal, not discovery. In LitmusChaos, the CNCF update notes pipeline integration through LitmusCTL, SDKs and Terraform, so experiments can be triggered and their results read programmatically. With Chaos Mesh, a pipeline step can apply the YAML, run the watcher script above, and delete the resource in a cleanup step that always runs, even on failure. The always-run cleanup is not optional. A cancelled pipeline that leaves a network delay in place is how a chaos experiment becomes an outage.
Gate on the right thing. A pass means the steady-state query stayed in tolerance, not that nothing happened. Capture the metric traces as pipeline artefacts, so a failing run can be diagnosed later without reproducing it.
Running a game day
A game day is a time-boxed exercise, usually two to three hours, in which a facilitator injects a pre-agreed set of faults while the on-call team responds as if it were real. The method that works is plain. Choose a scenario from a past incident or a known weak spot. Write the hypothesis and the abort criteria in advance. Name a facilitator who controls the injection and a scribe who records timestamps. Name an observer for each dashboard. Start with a quiet announcement so that stakeholders are not surprised.
Then run it in escalating stages. First verify steady state and alerting. Inject the first fault and measure time to detect, time to diagnose and time to mitigate separately, since they fail for different reasons. Detection fails when alerts are missing or noisy. Diagnosis fails when dashboards do not show the dependency. Mitigation fails when the runbook is stale or the person with permission is asleep. End with a blameless review that produces a short list of owned, dated actions. A game day with no action items either found a very robust system or was too gentle.
Many teams also use game days to rehearse failure of the chaos tooling itself: what do you do if the daemon is stuck and a delay is still applied? Practising kubectl delete of the fault resource, then confirming recovery on the node, builds the confidence to run experiments with less supervision later.
Open Source vs Managed Chaos Services: AWS FIS and Azure Chaos Studio
Managed services do not replace Chaos Mesh or LitmusChaos so much as wrap them with cloud identity, guard rails and infrastructure-level faults. Both major clouds can either inject Kubernetes faults themselves or call the open-source tools from inside a cloud-managed experiment, and the combination is often the strongest setup.
AWS Fault Injection Service (FIS). FIS offers native aws:eks:pod actions for standard Amazon EKS clusters: CPU, memory and I/O stress, pod delete, and network blackhole-port, latency and packet-loss. Per the AWS FIS documentation for EKS pod actions, FIS creates a pod in the target cluster that attaches an ephemeral container to the target pod to inject the fault, except for pod delete. The documented limits are worth reading closely: the network actions do not work on AWS Fargate or the bridge network mode, they need root in the ephemeral container, targets cannot be selected by ARN or tag, the pods must not set readOnlyRootFilesystem: true, EKS Anywhere and Hybrid Nodes clusters are unsupported, and the stated minimum EKS version is 1.30. Setup needs an IAM experiment role, a Kubernetes service account with a Role and RoleBinding in the target namespace, and an identity mapping from IAM to Kubernetes. Separately, AWS announced in July 2022 that FIS can run Chaos Mesh and Litmus experiments through the aws:eks:inject-kubernetes-custom-resource action. That lets one FIS experiment, for example, stress a pod with Chaos Mesh while terminating a share of nodes with native FIS actions. FIS stop conditions can be tied to CloudWatch alarms, which gives you the abort-on-SLO behaviour described earlier, although the sample EKS template in the pod-actions page sets the stop condition source to none, so you must add an alarm yourself.
Azure Chaos Studio. According to Microsoft’s tutorial on Chaos Mesh faults with the Azure CLI, Chaos Studio uses Chaos Mesh as the injection engine for AKS: Chaos Mesh is a service-direct target that you install first with Helm, you onboard the cluster as a target of type Microsoft-AzureKubernetesServiceChaosMesh, enable a capability per fault (for example PodChaos-2.1), and define the experiment fault with a Chaos Mesh jsonSpec. The experiment’s managed identity needs the AKS RBAC Cluster Admin and Cluster User roles on the cluster. Documented limitations include Linux node pools only, extra VNet injection configuration for private clusters, and allow-listing the ChaosStudio service tag if the API server restricts IP ranges. I have not verified the current full Azure fault library for non-Kubernetes targets, so check the fault library page for your region.
Choosing between them
| Situation | Open-source only | Managed service wrapping open source | Managed native faults |
|---|---|---|---|
| Single cloud, Kubernetes-centric | Good | Good | Good for node and infra faults |
| Need to fail a cloud dependency (zone, database, managed queue) | Poor | Best | Best |
| Multi-cloud or on-premises clusters | Best | Partial | Not available |
| Fine-grained network, I/O, time faults | Best | Best | Limited |
| Centralised approvals and audit | Build it | Built in | Built in |
| Cost and lock-in | Lowest | Medium | Medium to high |
The practical pattern: use the open-source engine for application-level faults you want in Git and CI, and use the cloud service when the fault is outside the cluster, such as an availability-zone impairment, or when governance requires a central approval and audit trail.
Trade-offs, Gotchas, and What Goes Wrong
Chaos tooling is itself a production risk, and the most common failures are self-inflicted. Treat the tool as privileged software that needs the same change control as a CNI plugin, because in effect it is one: its node agent runs with elevated privileges and manipulates networking and filesystems.
Unrecoverable faults. Network and I/O faults rely on a clean teardown. If the controller crashes, loses the node agent, or the node is partitioned during an experiment, a delay or filesystem fault may linger. Mitigate with short durations, which most fault types support natively, and with documented manual recovery steps tested in a game day.
Weak steady state. If the success metric is an average, a deep failure in one tenant disappears into it. If the metric has a five-minute window, a thirty-second brownout may not breach it. Fine-grained percentiles and per-route breakdowns reveal what averages hide, and synthetic probes exercise paths real traffic does not.
False confidence in staging. Staging rarely has production’s traffic mix, data volume, connection reuse, or noisy neighbours. A passing staging experiment shows that the mechanism works, not that production is safe. This is why the Principles of Chaos Engineering advocate production experiments, but doing so responsibly needs mature observability, narrow blast radius and organisational trust. Treat production chaos as a graduation, not a starting point.
Probe and kill semantics. A killed pod with gracePeriod: 0 tests crash recovery, while a normal rolling update tests the graceful path. They exercise different code. Likewise, a pod-failure fault leaves the pod object present but unavailable, which stresses readiness gating and endpoint removal in ways a kill does not.
Interaction with autoscalers and operators. Horizontal and cluster autoscalers can mask a fault by adding capacity, and a stateful operator can react to a killed primary by triggering a failover that is more disruptive than the fault. Know which reconciliation loops are running before you inject.
Stateful systems. Killing a database replica may be a valid test, but killing it during a rebalance or backup can corrupt a recovery. Schedule these around maintenance windows, verify backups first, and keep the fault count low.
Alert fatigue and cost. Experiments will trigger alerts. Either silence the specific alerts with an expiry and a tag, or, better, make the alert firing part of the hypothesis, because an alert that does not fire is a finding. Remember also the cost of the installation itself: daemon sets on every node and a dashboard are small but non-zero.
Version and compatibility drift. Chaos Mesh’s own documentation says support for newer Kubernetes versions may be theoretical rather than tested, and LitmusChaos moves monthly. Pin versions and test upgrades.
Practical Recommendations
Start small and build a ladder. Your first month should produce one working experiment, one abort mechanism and one finding that changed something. Resist the urge to install the full catalogue; a single well-understood fault with a good hypothesis beats thirty unexamined ones.
Choose a tool by team shape. If you have a platform team comfortable with Prometheus, Argo or Tekton, start with Chaos Mesh and write the judge yourself. If you want a self-service product with probes and scores, evaluate LitmusChaos and its ChaosCenter. If your faults live in the cloud rather than the cluster, add AWS FIS or Azure Chaos Studio and let them orchestrate the open-source engine.
Make the safety controls non-negotiable before the first run, and treat them as requirements rather than refinements. Then expand along one axis at a time: more replicas, then longer duration, then a harsher fault, then a more realistic environment. Keep the experiments in Git, review them like code, and link each one to the hypothesis and the incident or risk that justified it.
A short checklist to adopt:
- Steady state defined as user-visible SLIs, with a numeric tolerance.
- Namespace filter on, and chaos allowed only in annotated namespaces.
- Every fault has a duration and a bounded
mode. - An independent abort watcher, tested by deliberately breaching it once.
- Always-run cleanup in pipelines.
- Chaos controller and monitoring excluded from fault selectors.
- Namespace-scoped RBAC for product teams, cluster scope for the platform team only.
- A game day each quarter, with owned, dated action items.
- Pinned tool versions, upgrade tested in a non-production cluster.
If your platform also runs autonomous operations tooling, our analysis of AIOps incident response and agentic SRE architecture shows how to feed chaos findings into automated remediation. For edge and device-facing clusters, see how Azure IoT Akri exposes devices as Kubernetes resources, a useful case where hardware-dependent failures need their own experiments.
Frequently Asked Questions
What is chaos engineering in Kubernetes?
Chaos engineering in Kubernetes is the practice of deliberately injecting failures, such as pod kills, network latency or node loss, into a cluster to test whether a defined steady state holds. Experiments are expressed as custom resources, run with a limited blast radius, and aborted automatically if service-level indicators breach tolerance. The goal is to find hidden weaknesses in retries, timeouts, probes and failover before real incidents do, and to turn each finding into a fix.
Chaos Mesh or LitmusChaos: which should I use?
Choose Chaos Mesh if you want fine-grained fault primitives (pod, network, I/O, time, stress) that you script into your own pipeline and judge with your own metrics. Choose LitmusChaos if you want a control plane, a hub of experiments and built-in probes that give each run a pass or fail verdict. Both are CNCF incubating projects. Many teams trial each against one service and keep whichever fits their workflow.
Is it safe to run chaos experiments in production?
It can be, but only after you have mature observability, a narrow blast radius and automatic abort conditions. Start in staging, enable namespace filtering, use short durations and bounded modes, and restrict who can create fault resources. Move to production one service at a time, during staffed hours, with a rollback plan. Production experiments reveal real traffic behaviour, but they are a graduation step and not a starting point.
What is a steady-state hypothesis?
A steady-state hypothesis is a falsifiable prediction that a measurable, user-visible output stays within tolerance while a fault is injected. For example, the checkout success ratio stays above 99.5 percent while one replica is killed. You define steady state first, predict it holds in both control and experimental groups, inject the fault, and try to disprove the prediction. A disproved hypothesis is a valuable finding, not a failure.
How do I stop a chaos experiment that is going wrong?
Delete the fault custom resource. Both Chaos Mesh and LitmusChaos reverse the injection when the object is removed or its duration expires. Automate this with an independent watcher that queries your SLO metrics and deletes the resource on breach, and add an always-run cleanup step in CI. Rehearse manual recovery on a game day, including the case where the node agent is unreachable and the fault lingers.
How often should we run game days?
Quarterly is a sensible starting rhythm for a service with an on-call rotation, with automated chaos checks running on every release or nightly in between. Run an extra game day after a major architecture change, a significant incident, or onboarding of new on-call engineers. Keep each exercise to two or three hours, with a written hypothesis, a facilitator, a scribe and a blameless review that yields owned action items.
Further Reading
- AIOps incident response architecture and agentic SRE in 2026
- Azure IoT Akri: a Kubernetes resource interface for devices
- Cilium vs Calico for industrial Kubernetes networking: an ADR
- Cilium vs Calico for industrial Kubernetes networking, second look
- Principles of Chaos Engineering for the canonical method
- Chaos Mesh documentation and LitmusChaos documentation for fault references per version
- AWS FIS EKS pod actions and Azure Chaos Studio with Chaos Mesh on AKS
By Riju — about
