Cloud Cost Reduction
CIOREVIEW >> Digital Twin >> NEWS

Lolli & Pops

Norman Paulsen, VP IT & Digital

Cloud Cost Reduction

Norman Paulsen, VP IT & Digital
Norman Paulsen, VP IT & Digital, Lolli & Pops

Whether it is Basecamp’s impressive 3.2 million USD cloud bill or Dropbox’s migration out of the cloud, there is a lot of buzz right now around the high cost of cloud computing. Go back just a few years and every IT leader was moved to or moving to the cloud. Now, you cannot avoid stories of run away bills or murmurs of moving back to on-prem servers. Snap cut computing costs by 65 percent and Dropbox by 74.6 million USD over two years by moving away from the cloud. When companies put up these big numbers, you cannot help but take notice and wonder: is my cloud bill too high? I have seen a good number of cloud implementations and I can say with some clarity, absolutely.

When I joined Lolli & Pops, a mid-sized candy retailer, they had the same problem. As my first foray into retail, it was interesting to see they suffered from the same problems as the IT industry when it came to cloud costs. Through some performance tuning and optimizations, I was able to cut Lolli & Pops bill by 85 percent in under a year. We have even further savings that will be realized in the coming months getting us over a 90 percent reduction in costs. I am here to tell you that you do not need to build your own servers to save on your cloud bill. I will detail out a few of the quick wins that can drastically reduce your cloud infrastructure costs quickly.

Knowledge is power.

Become familiar with your bill. Understand how to slice and dice your daily cloud costs. You will be surprised where that spend is going. Important slices to know are spend by usage type, spend by instance type and spend by region. Usage type breaks down your spend into finely grained buckets including instance types, volume storage, gateway hours, etc. The instance type report shows spend by VM size. The spend by region report shows what zones your spend is occurring in. All of these reports help identify systems you are paying for that you are not using.

We found several stagnant systems and resources we were being billed for but were not in use. Cleaning these up is a quick way to cut costs. We found systems running in regions that were not in use and high usage spend on volume storage for archives no one had looked at in over a decade. Understanding your cloud spend at a granular level is the first step to cutting costs.

Rightsizing

The next step was correctly sizing some of our servers. The instance report allowed us to identify our most costly virtual machines. Armed with that information, we knew where to focus first. We had a small subset of servers accounting for most of our cloud spend. Not wanting to break our existing infrastructure, monitoring was the next critical step. We installed monitoring and log ingestion tools on these servers to pull all key metrics and to track what was occurring on these servers. One tip, be sure to check for cron jobs and open ports. That will help scope out the server’s usage. Based on our findings, we detailed out what was running on each machine, how it was being utilized and what resources it was consuming. We had instances running at under 0.1 percent CPU utilization and targeted them first. We were able to quickly pause those instances, resize them and bring them back online. We moved CPU utilization closer to the 5 percent utilization, which cut our spend significantly.

 You will be surprised at what you can save in cloud cost without moving away from the cloud 

Our 5 percent CPU utilization target was arbitrary and really depends on peaks in usage. Next, it is important to have a production like test environment to try different instance types to find the best utilization of resources that doesn’t impact performance. If you have a job running only once a day or once a week, spiking CPU significantly, you may consider moving it to a lambda function to save further. Instead of paying for the idle CPU throughout the day, you will only incur compute costs when the job runs.

Reducing Risk

It is always concerning shutting off a server, especially when you are not entirely sure of everything it is doing. However, in most cases the savings always out ways the risk. There are a few additional things you can do to further reduce the risk of an outage. Instead of deleting an instance, you can shut it down instead. Then, if you determine it is needed, you can simply start it back up. If that seems too risky, you can just limit the instance’s access via its firewall so the instance is never shut down, only temporarily unreachable. For our instances, after I felt we had a good handle that it was not in use, we would spin down the server for a week before terminating it.

Cloud Cleanup

Based on most of the implementations I have seen over the years, organizations have plenty of room to cut spend in their cloud environment without reducing quality. In fact, these types of exercises typically uncover some poor design patterns that ultimately allow for a better cloud architecture. In Lolli & Pops’ case, we had four separate servers running on one virtual machine: Payara, Tomcat, Nifi and MySQL. We were able to both reduce costs and isolate these four servers for better uptime and redundancy. I would like to note, you should never run a database on a standalone VM. Use a redundant database cloud service like RDS or GCS instead.

While you may not see an 85 percent cloud cost reduction, like Lolli & Pops, you will be surprised what you can save without moving away from the cloud. If the tasks seems too daunting, hiring an outside cloud performance agency is always an option. Good luck on the next leg in your cloud journey.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.