AI and Digital Transformation in an Established R&D Enterprise
CIOREVIEW >> Big Data >> NEWS

Procter & Gamble [NYSE: PG]

Kelly L. Anderson, Senior Director, PS Data Science, Corporate IT

AI and Digital Transformation in an Established R&D Enterprise

Kelly L. Anderson, Senior Director, PS Data Science, Corporate IT

One of the hottest trends right now in mid- to large- corporates is attempting to leverage the existing digital assets of the enterprise and to create an ‘AI Factory’. This AI Factory (introduced by Marco Iansiti and Karim R. Lakhani in their book “Competing in the Age of AI”) aims to create a foundational layer of data that can be activated in many experiments and monetized with almost zero marginal costs. The idea is brilliant, and there are many examples of where it has been successfully implemented, especially for new ‘digital-first’ companies as they scale. However, most established mid- to large- size enterprises that have many legacy systems and a culture of doing things a certain way, are still far from living the AI-Factory life. How can they close the gap?

What is needed:

• A very clear business value hypotheses (i.e. the ‘so what’ of this effort). What is the most impactful business task that would benefit from our valued and scarce expert data science resources working on? What will provide the most sustained ROI and justify continued investments in building the foundations to the AI Factory.

• A high performing team that is fully sponsored and empowered to implement the changes that may disrupt current business operations and work processes. A team united in their vision to deliver this AI Factory and proven in their ability to get the job done.

• An Urgency to get started. Perhaps first with well structured and pre-labeled tabular data that has clear business value, has existing work processes around it, and can be the baseline for the AI-factory.

• Then, for unstructured data, new modalities of data, or new business ideas… first use unsupervised methods or semi-supervised methods, and then when the problem definition and value creation is well defined, a sustained effort to build the data capture system, annotate the data and build the fully supervised methods for the ultimate business deliverable to be put into production with a robust data pipeline.

“The idea is brilliant, and there are many examples of where it has been successfully implemented, especially for new ‘digital-first’ companies as they scale.”

• Patience to see (super-) human-level performance. Model improvement progress is rarely linear, but happens in bursts and spurts as the work processes mature, as data driven model building is optimized, and evolves as gaps in training data are filled by intentionally collecting and annotating the most useful data to improve the model’s performance vs. simply throwing more data at the problem.

• An operations team that can exploit the models and fully operationalize the AI Factory to serve ‘living models,’ monitor their performance and usage.

• Feedback from users and the operations teams, and investments for ongoing maintenance, model updates, new data feeds, A/B testing. Open lines of communications from in-market support to upstream R&D to ensure new requirements are captured and actioned.

• Lifecycle management. For data, systems, and models. Fully documented.

On the deep technical side, key questions business leaders may include: how do we better utilize our existing assets and tap into the emerging AI capabilities such as; large sparse tabular dataset modeling, graph neural networks to connect ‘puddles’ of data and models, transformers for natural language understanding and computer vision, generative / diffusion models that are leading to ‘creative computing’ applications and create business value.

The reality is that much of the data in the enterprise is not structured, not accessible, not documented sufficiently, and was likely not collected with the intent to be consumed by an algorithm. This doesn’t mean it should all be ignored, but know that additional effort will be needed to refactor this data so that it might be useful. Likely costs include re-basing the data to be fed to an algorithm, annotating data for new features of interest to the business question to be answered, and investments in re-training the workforce to work in a new way to better feed the algorithms and to integrate results from models across modalities or disciplines to gain insights.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.