Predictive Analytics still rules, but Context is King!
CIOREVIEW >> Predictive Analytics >> NEWS

Associate Director of Data and Analytics at Ocean Spray Cranberries

Paul Cavacas

Predictive Analytics still rules, but Context is King!

Paul Cavacas

With the explosion of Gen AI the AI/ML/Gen AI arena it has been getting all the press and focus, while Gen AI provide tremendous potential and can be a once in a generation transformational shift, one should not underestimate more traditional AI and ML in today’s environment.  Gen AI works extremely well for general text, and to a lesser extent image, video, and audio generation, but when presented in a more typical business environment where hallucinations and false positives can cost the company thousands or millions of dollars you may be better suited to switch some focus back to the more traditional predictive analytics.  Having said that, time should be invested in Gen AI as well, but don’t overlook even something as simple as linear regression to provide more predictable value.

Predictive Analytics has become so common place that even nontechnical people expect to just have it available, whether you are talking about a recommendation engine or demand forecast.  All these use cases have just become embedded into our everyday way of thinking and while the process and tools needed for this are well established and readily available to all, you cannot overlook the need for good solid data and a thorough understanding of what the data means and external factors that play into the predictions.

Taking a real-world example for Ocean Spray we generate a regional crop forecast for the cranberry growers that are part of the cooperative.  This forecast is used to plan out the supply for the entire cooperative and is a pivotal part of the business.  Over the past several years this predictive forecast, done using traditional ML methods, considering things like history, weather, and various other measurements has performed well and been highly accurate. 

This past year the prediction was a lot less accurate than other years, and nobody really knows the underlying reason why the crop was considerably higher than any prediction.  One thought about why this may have happened is that there has been an invasive plant species that has been growing near the bogs over the past decade or so and the bees that are brought in to pollenate the bogs have been pollinating this other plant instead of the berries in the bog.  This past year having been extremely rainy has stunted the growth of this other plant, so the bees could pollinate the berries as planned.

The moral of the above story is that even in established models that have been used in the past there is always more to learn about the context around the data and what it represents.  The intricacies of data and the relationships between the different variables are far more complicated than you might originally think, so continually reevaluating results and validity of models is paramount.

The best results will always be achieved when the Data Scientist understands these intricacies and is able to codify them into different features and rules, which may need some ingenious solutions to achieve.  The ease of use and power of the existing tools will mask this necessity and allow even the beginner Data Scientist to achieve good results, but to really advance into the greatness realm you need to understand data, its context and have the tools in your toolbelt to be able to make the most of them, whether that is Python coding, SQL queries, or other most exotic tools, having a strong foundation will get you a lot farther then the tools and models themselves.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.