Reliability · Incident commander
Checkout is melting under a retry storm
A slow payment dependency turns ordinary retries into a cascading failure across checkout.
Work through realistic incidents with incomplete evidence and competing constraints. Each exercise asks you to diagnose the failure, stabilize the system, recover safely, and prevent recurrence.
Showing 8 scenarios
Reliability · Incident commander
A slow payment dependency turns ordinary retries into a cascading failure across checkout.
Consistency · Service owner
Failover restores availability, but asynchronous replicas expose older account state.
Consensus · On-call engineer
Uneven latency and pauses create election churn without an obvious node failure.
Partitioning · Platform engineer
A technically balanced keyspace collapses when one key becomes globally hot.
Messaging · Application architect
At-least-once delivery meets a non-idempotent warehouse side effect.
Architecture · Principal engineer
Active-active inventory remains available through a partition and violates uniqueness.
Coordination · Distributed systems engineer
Time-based ownership outlives its safety assumptions after virtualization pauses and clock correction.
Messaging · Data platform owner
Consumer lag, oversized batches, and unbounded buffering turn degradation into data loss risk.