Architecture

7 Chaos Engineering Examples That Build Resilience

A payment API can be healthy right up until its database slows down, its cache returns stale data, and one upstream service begins timing out at the same moment. Chaos engineering examples give teams a controlled way to find out what happens next – before customers, revenue, or on-call engineers pay the price.

Chaos engineering is not random breakage in production. It is a disciplined resilience practice: form a hypothesis about how a system should behave under stress, introduce a limited fault, observe the result, and improve the design or operations where reality differs from expectations. For teams running distributed applications, that feedback loop is far more useful than assuming redundancy alone will handle every failure.

What Makes a Chaos Experiment Useful

A useful experiment starts with a steady-state definition. This is the measurable behavior that indicates the service is working normally, such as successful checkout completion, acceptable API latency, message-processing throughput, or error rates below an agreed threshold.

The experiment then changes one condition while keeping scope deliberately small. A team might affect a single Kubernetes pod, one availability zone, a noncritical dependency, or a small percentage of requests. Guardrails matter just as much as the fault itself. Define abort conditions, assign an observer, choose a low-risk time window, and make sure the team can stop the experiment immediately.

The goal is not to prove that a component can fail. Every component eventually can. The goal is to verify that the system fails in an expected, recoverable, and visible way.

7 Chaos Engineering Examples for Modern Systems

1. Terminate an Application Instance

One of the most approachable chaos engineering examples is intentionally terminating an application instance. In a containerized environment, delete one pod. In a virtual machine environment, stop one instance. The hypothesis may be that the load balancer will route traffic to healthy replicas and autoscaling will replace the lost capacity without users noticing.

This test exposes issues that are easy to miss in architecture diagrams. A service may have multiple replicas but still keep session state in local memory. Startup probes may allow traffic before the application is ready. Replacement instances may pull large images or wait on migrations, extending recovery beyond the team’s assumptions.

Measure customer-facing latency, failed requests, replacement time, and the behavior of any in-flight jobs. If one terminated instance causes a visible outage, the answer may be more replicas, but it could also be stateless session handling, better readiness checks, or faster deployment artifacts.

2. Add Latency to a Critical Dependency

Most distributed-system failures begin as slowness rather than a clean outage. Add 500 milliseconds, two seconds, or a carefully selected delay to calls between an application and a database, identity provider, payment processor, or internal API.

The hypothesis should be specific: requests should time out within a defined budget, retries should remain bounded, and the application should return a useful fallback or error instead of tying up every worker thread. That distinction matters. Retrying a slow request can improve reliability when the failure is transient, but aggressive retries can amplify load and turn a partial problem into a broader outage.

This experiment often reveals missing timeout values. Without explicit connection and request timeouts, a dependency can consume application threads until the service can no longer handle otherwise healthy traffic. It also tests whether circuit breakers, bulkheads, and fallback responses work as designed.

3. Simulate a Database Failover

Databases are frequently engineered for high availability, but application behavior during failover deserves separate testing. Simulate the loss of a primary database node, trigger a planned failover in a safe environment, or temporarily block connections to a replica.

A well-designed system may experience a short connection interruption, retry idempotent operations, and restore service after the new primary is available. A less prepared system may generate duplicate orders, lose writes queued only in memory, or hold database connections that never recover.

Focus on consistency as well as uptime. For example, an e-commerce service can safely retry a read, while retrying a payment capture requires an idempotency key and clear transaction state. This is where resilience engineering meets business logic. Technical recovery is not enough if the recovery path creates customer or financial errors.

4. Exhaust a Shared Resource

Shared resources fail differently from individual instances. Introduce controlled CPU pressure, memory pressure, disk exhaustion, file-descriptor exhaustion, or connection-pool saturation in one part of the system. The expected response might be that nonessential workloads are shed first while core user actions remain available.

Connection pools are particularly valuable to test. A service can look healthy under normal load while a slow downstream dependency gradually consumes all available connections. Once the pool is exhausted, new requests queue, latency climbs, and the impact spreads to callers.

The corrective action is not always to increase a limit. Larger pools can place more pressure on an already struggling dependency. Teams may need request limits, queue backpressure, workload prioritization, or separate pools for critical and background traffic.

5. Break DNS or Service Discovery

Applications often depend on DNS, service meshes, registries, and load-balancing layers that are invisible during feature development. Simulate failed DNS lookups, stale service records, or an unreachable service-discovery endpoint.

The hypothesis could be that existing connections continue serving requests, DNS caching prevents a short interruption from becoming an outage, and failures appear clearly in telemetry. In practice, some runtimes cache records longer than expected, while others resolve names so frequently that a DNS problem becomes an application problem immediately.

This example is especially useful for cloud-native systems where workloads are short-lived and service endpoints change often. Test both the initial connection path and the behavior of long-running instances after records change. A resilience plan that works during deployment may behave differently during a regional networking incident.

6. Introduce Message Queue Delays and Duplicates

Event-driven services need experiments tailored to asynchronous failure modes. Delay message delivery, make a consumer unavailable, inject duplicate events, or send messages out of order. These conditions are normal possibilities in distributed messaging, even with reliable managed platforms.

A good outcome is not necessarily instant processing. It may be a growing but observable backlog, autoscaling consumers, preserved ordering where required, and no duplicated side effects. For an order-processing workflow, the same event delivered twice should not create two shipments or send two refund requests.

Test dead-letter queues as part of this work. A dead-letter queue is only useful if teams can identify why messages arrived there, replay them safely, and avoid immediately recreating the original failure. Alerting on backlog age is often more meaningful than alerting only on queue depth, because a large queue can be acceptable during a planned traffic spike while aging messages indicate delayed customer outcomes.

7. Disable Observability Inputs

A system can recover from a fault while the team remains unable to explain what happened. Temporarily block a metrics exporter, drop a tracing collector connection, or disable logs from one service in a controlled environment. This is a chaos experiment for operational visibility.

The objective is to confirm that monitoring has useful redundancy and that on-call responders can still identify impact. If a dashboard reports green because it relies on one missing telemetry pipeline, the dashboard is measuring data delivery rather than customer experience.

Pair infrastructure signals with service-level indicators such as successful transactions, latency percentiles, and completed workflows. Synthetic checks and business metrics can continue to show a problem when a specific host agent or log pipeline fails. Observability should help teams make decisions under pressure, not create a second incident to investigate.

How to Run Experiments Without Creating an Avoidable Incident

Start in a staging environment that resembles production, but do not stop there. Staging rarely has the same traffic patterns, dependency behavior, scale, or operational complexity. After validating the experiment, use a narrow production scope only when the system has clear rollback options and the potential impact is understood.

Write the hypothesis before running the test. For example: “If one checkout pod fails, completed checkout rate will remain above 99.9%, p95 latency will stay below 800 milliseconds, and the platform will replace the pod within three minutes.” A statement like this turns a vague resilience exercise into an engineering test with a pass or fail result.

Teams should also distinguish between experiments that test component recovery and those that test organizational response. An automated pod replacement may pass technically, while unclear alerts and missing runbooks still leave the on-call engineer guessing. Both findings are valuable, but they call for different improvements.

Choose experiments based on real risk. Incident history, architecture reviews, capacity limits, and dependency maps are better starting points than copying fashionable failure scenarios. A small team operating one API may gain more from testing database timeouts than from simulating a full regional outage. A multi-region platform with strict availability commitments may need both.

The best chaos engineering program becomes part of ordinary engineering work. Each experiment turns an assumption into evidence, and each finding becomes a practical change to code, infrastructure, monitoring, or operational documentation. Start with one failure your team would hate to learn about from a customer, then make it safe to learn about it yourself.

Related Articles

Back to top button