The Data Foundation for Production-Ready AI
CIOREVIEW >> Artificial Intelligence >> NEWS

This article is part of CIOReview's Innovation Insights series featuring expert contributions nominated by our subscribers and reviewed by our editorial team.

The Data Foundation for Production-Ready AI

Stephen Godfrey, Co-Founder, Numantic Solutions

Data Strategy Authority

Editor’s Note: CIOs pursuing AI at scale need to treat data quality, measurement and system flexibility as core infrastructure rather than secondary implementation concerns. Stephen Godfrey’s perspective matters because durable data assets can strengthen model performance, support responsible evaluation and give organizations a more resilient foundation for production AI.

Building the Data Advantage before the AI Advantage

Prior to forming Numantic Solutions, our team worked on building and deploying production-grade machine learning and AI solutions for large organizations with well-resourced and mature data science teams. Even in those situations, we were keenly aware that the quality of model outputs was highly dependent on input data. 

Now that LLMs and other AI tools are widely accessible, the importance of input data has become even more critical. Datasets are the secret sauce in many successful AI production deployments because they drive the relevance and quality of outputs, and they can be critical in building test harnesses that measure AI performance which is often very difficult.

This led us to form a consultancy primarily focusing on smaller, less-resource organizations. We think AI can help level the playing field by assisting small organizations scale, but they need to be builders and combine data, analytics and technology to construct differentiated solutions. We believe we can help.

It’s always been about the data, but now so more than ever. The good news is that even small organizations can build up high-quality, unique datasets supporting production AI solutions with a relatively modest investment. And building such datasets will probably provide the highest level of return on investment with their AI investments

Start with the Value, Then Build the Data

Leveraging data starts with a vision or strategy of how data can make the organization more effective, efficient and in doing so further its mission. Developing such a vision begins with an objective perspective on how the organization is delivering value and where in operational procedures could higher quality or more readily available data accelerate the delivering of services. It’s important to note that the data vision doesn’t need to be fully baked, and it’s likely to evolve over time. As data solutions are developed additional opportunities will be identified providing the foundation for a virtuous cycle of continual improvement. 

With a data vision, organizations are more likely to allocate resources to build solutions that deliver actionable insights and deliver value to the organization. The data strategy provides a line of sight into the investment’s benefits and helps establish metrics that ensure their data science approach is delivering value.

  ​The good news is that even small organizations can build up high-quality, unique datasets supporting production AI solutions with a relatively modest investment.   

The good news is that even a modest investment in building data solutions and their underlying datasets can have a huge impact on the quality and effectiveness of insights. 

AI Needs a Scorecard, Not Just a Use Case

Any technological solution addressing complex business problems requires performance measure frameworks especially if they incorporate automated decisioning. This is true even if the solution’s pipeline or process has human-in-the-loop checkpoints. 

Measuring modern AI tools can be particularly challenging since they can handle a wide range of tasks, often have unstructured inputs or outputs such as free-form text and frequently sound convincing even when hallucinating. In addition, these solutions need to be assessed across multiple dimensions including safety and responsibility.

Despite challenges, frameworks can be built to assess AI tools. In designing such measurement solutions for production solutions, we typically rely on four design principles:

First, AI components should have well-defined inputs and expected outputs. Even in cases like chatbots in which inputs may not be fully definable since users can theoretically enter any question or comment, expectations of outputs to these ‘random’ inputs should be developed. 

Second, measurement requires curated data which is often used as both AI inputs and AI benchmarks. 

Third, datasets should be designed to expand over time and in many cases should incorporate information captured once the system goes live.

Fourth, performance assessments are often optimized by including evaluations from both human experts as well as other independent AI tools. And, these assessments should cover multiple dimensions including accuracy but also safety, objectiveness and fairness.

Building the Data Assets That Outlast AI Models

AI is driving demand for data, and many data providers are reporting a surge in demand. That is understandable since AI tools can handle large amounts of relatively unstructured data and effective and differentiated AI solutions often rely on unique datasets. In addition, we see many builders realize that their curated datasets are their most defensible assets. In an era of low-cost access to foundational models and open-source tools, the technological foundations of most solutions are similar. It’s only through unique datasets that they differ. So, we expect a continued focus on data.

But data is not the only factor rapidly driving change. Without question, the speed of AI development is astounding. It’s not just that there’s a new foundational or open-source LLM released almost every week, the ecosystem of supporting components is also quickly evolving. This means that most production AI solutions will seem out-of-date in a year or two. However, that is unlikely to happen to high-quality data.

In our view, builders should design for future flexibility. Data should be carefully curated and designed to continually grow. AI systems should be modular and built in ways that allow for easy upgrades or replacement.

Pairing AI Fluency with Domain Expertise

It is both an inviting and intimidating time to build a data science career. It’s inviting because the international attention on AI is driving demand for solutions and because LLMs make it much easier to quickly build solutions. Everyone should feel empowered and encouraged to use a coding agent to start building a solution of interest. It’s intimidating because AI may replace workers, AI tools are rapidly evolving and a vibe-coded process may go rogue or fail in production. 

For these reasons, my advice would be to stay grounded and become a subject matter expert in the business or other discipline problem being addressed. View AI as a tool that you direct and keep it focused on the right problems. This will help ensure solutions are adding value and are working as expected. 

MORE FROM INNOVATION INSIGHTS

The Data Foundation for Production-Ready AI
Numantic Solutions
Stephen Godfrey, Co-Founder, Numantic Solutions
What does AI-ready even mean?
DAS42
Susan Payne Cook, CEO
Cloud Hosting in 2026: What's  Actually Going On
All in IT
Dave Primley, Solutions Advisor
One Plate, One Platform: The Future of Smart Parking Management
Mobile Smart City Corp
Luis Garma, Founder and Chairman

EXPLORE OUR KNOWLEDGE NETWORK



The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.