Speed Vs Stability - An Interesting Paradox
CIOREVIEW >> Odoo >> NEWS

Northwestern Mutual

Lalit Arora, Senior Director of Engineering, Infrastructure and Cloud Services

Speed Vs Stability - An Interesting Paradox

Lalit Arora, Senior Director of Engineering, Infrastructure and Cloud Services
Lalit Arora, Senior Director of Engineering, Infrastructure and Cloud Services, Northwestern Mutual

To be speedy or to be stable, that is the question. Or is it?

Developers strive for speed of delivery (to meet their customer demands) and sysadmins need stability (to meet their customer demands). Interestingly, their customers are the same, but why do developers and sysadmins chase different goals? Quite simple. They are incentivized for different goals.

Developers are incentivized to rapidly deliver features (speed/velocity), while sysadmins are incentivized for stability. Are speed and stability mutually exclusive? Can both teams achieve their respective goals?

Like many things in life, it’s not exactly black and white, but let’s try to understand these seemingly competing objectives with the help of a few questions:

Are we Chasing the Right Availability Target?

Let us start by determining what the SLA/SLOs are for the services (app and infra) we are responsible for. These are most likely already set and stored someplace. The hard part is determining if they are appropriate, realistic, and if the business/product owners even know what the SLAs mean.

In my recent experience of leading the disaster/disruption recovery (DR) program, we found that many apps and infra were assigned the recovery time objective (RTO) of zero or one hour, which means they are expected to recover from a disruption either immediately or within an hour. A lot of effort and capital would be required to meet those expectation. After talking to business owners, we realized they did not understand the true cost of the RTOs they were assigning and were quite willing to adjust them.

This does not mean that the current service SLAs are wrong, but it is certainly worth exploring. 99.99 percent uptime (4 nines) means only 52 minutes of downtime is allowed per year, whereas 99.9 percent uptime (3 nines) permits 8 hours and 45 minutes of downtime. Let us figure out how we can use this difference in downtime.

What is Our Error Budget?

Atlassian explains the error budget quite well— An error budget is the maximum amount of time that a technical system can fail without contractual consequences.

Regardless of role (Developer, Operations, or DevOps), we should know what our application/platform/service's error budget is. I will admit that while I knew the availability requirements, most of the time, I never knew my error budget. Understanding the error budget can help us effectively leverage that budget to innovate (speed) and take risks (stability).

Is There an Effective Change Enablement Process?

A common misconception among developers is that just because they provision and operate their application’s infrastructure, they do not need a change process. The truth is most organizations, are highly connected, and our work has the potential to impact other applications, services, business areas, etc.

We need to shift our mindset from change control to change enablement. Instead of restricting or controlling changes, we want to maximize the successful changes while ensuring risks are properly assessed, authorized, and managed before and during change windows.

How Quickly Can We Recover From an Outage?

Unless we are in the business of manufacturing pacemakers, we can expect our systems to break. Nevertheless, we should still focus on making our systems more reliable (recover quickly from an outage). And not just quickly, but in a methodical way. Is there a hero who saves the day every time there is an incident, or is there a documented playbook that includes well thought-out and anticipated outage scenarios with steps to leverage to recover from an incident? This playbook should not be “Build Once Forget Always” (BOFA), rather a living artifact that is updated with learnings from every major incident.

These are some of the ways developers and sysadmins can work together to achieve both goals (speed and stability). If asking these questions becomes part of our culture, we will be able to sustain the system stability we need to support the rapid development we want.

What has helped you in your mission to deliver faster yet meet availability/stability requirements?

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.