Hanging out under the same umbrella as chaos engineering, resilience engineering is a way of building your systems to fail. Let’s take a look.
What is resilience engineering?
Let’s start with resilience—the ability to keep on keeping on in the face of failure. In the words of Bob Dylan, “There’s no success like failure, and failure is no success at all.” And in my own, “Failure sucks.” In terms of technology and IT systems,
Resilience is a system’s ability to recover from a fault and maintain persistency of service dependability in the face of faults.
Resilience engineering, then, starts from accepting the reality that failures happen, and, through engineering, builds a way for the system to continue despite those failures. Good resilience engineering produces a system that can adapt. Resiliency can be built into any system, and it offers a lens to look at critical areas like cybersecurity and operations.
Here are some examples:
- If the system adapts by acquiring servers from a different region when all the servers in its zone fail, it has been successfully engineered.
- If the system adapts by taking the next best CPU when the cloud provider cancels providing the present CPU, the system has been successfully engineered.
- If the system fails to scale its number of servers when, suddenly, its number of users skyrockets, then the system has not been successfully engineered.
How adaptability supports resilience
Adaptability is the defining trait of resilience. If you are to provide a SaaS product, and your systems go down, there is no product.
Humans have long been the primary agent in making systems adapt. It has been people who are on-the-ready to investigate and get the software back up and running as quickly as possible—to make a system resilient to failure. Human attention has been required to ensure system resiliency.
But human labor is old-school in the age of software. How Complex Systems Fail by Richard Cook is a short document that covers common ways that systems fail. 50% of the choices have to do with human error or the necessity of human intervention.
Software, now, is being designed to help make systems adapt. With the dawn of cloud computing, and infrastructure parts like containers and Kubernetes orchestration, software is doing the work instead of people.