Assuring Data Quality is the Best Way to Improve Organizational Decisions
CIOREVIEW >> Quality Assurance >> NEWS

VCU/Data Blueprint

Peter Aiken, Associate Professor of Information Systems, Founding Director

Assuring Data Quality is the Best Way to Improve Organizational Decisions

Peter Aiken, Associate Professor of Information Systems, Founding Director
Peter Aiken, Associate Professor of Information Systems, Founding Director, VCU/Data Blueprint

A voluminous study recently documented billions of dollars in organizational costs as a result of poor quality data. These costs usually manifest as a result of low quality data used by smart individuals who use it to make decisions. It is a problem of garbage in - garbage out. Let me relate a specific example that I recently encountered at a large logistics firm.

“Data quality engineering must achieve a more complete picture and facilitate cross boundary communications” 

My team discovered a room full of 100 firm associates who worked full time correcting ALL bills as a prerequisite step to transmitting them to customers. Every bill contained large amounts of incorrect data ranging from the date of service to the type of service to the price charged by the firm and resulting in a 30 day delay on all receivables going out the door! When I pointed out that this was unnecessary expense and that it could be corrected with a relative easy data quality assurance initiative, the response was astounding by in context understandable. The manager in charge informed me that, since the firm had just experienced the best quarter of the year in its history, the thought was that perhaps the size of team should be doubled since it was clearly producing good results.  (A quick side trip to the CFO easily produced an ROI based on increasing cash flow on $9 million annually by 30 days and resulted in a more mature approach and hundreds of millions in additional revenue from the on-time receivables.)

Stay ahead of the industry with exclusive feature stories on the top companies, expert insights and the latest news delivered straight to your inbox. Subscribe today.

Many in businesses are still dealing with these challenges from various stove piped perspectives, approaching data quality problems in the same way that the blind men approached the elephant - people tend to see only the data that is in front of them. Little cooperation exists across organizational divisions, departments, even workgroups -just as the blind men were unable to convey their impressions about the elephant and collectively recognize the entire entity. In order to be effective, data quality engineering must achieve a more complete picture and facilitate cross boundary communications.

After dealing with data quality problems for more than 30 years, I have formed two strong opinions:

First, prevention is more cost effective than treating the symptoms. It should be obvious that correcting the data quality problems will be less expensive than fixing them forever.

Second, data quality problems are more unique than being similar. This prevents the resolution of these challenges from following programmatic solution development practices and it mandates the development of specialized data quality engineering specialists within organizations.

Data quality is now widely acknowledged as a major source of corporate risk, because of the above mentioned “garbage in garbage out” problem. After seeing the various response patterns repeated, I became aware that data quality was also a socio-technical discipline as evidenced by the need to involve the logistic company’s CFO. Our understanding of data assurance has resulted in a new understanding of our approach to data quality assurance (shown below).

Since the development of formalized data reverse engineering and the invention of data profiling, our collective data quality tool kit has matured considerably. A multitude of products are now available to help out with various analyses and tasks. The most common problem now facing organizations is the wide spread perception that tools alone will accomplish data quality improvements and that purchase of a data quality package will entirely solve data quality problems. As we continue to learn more about data quality, solutions engineering, and related issues, one thing will continue to remain clear: the best data quality engineering solutions will continue to be a combination of selected tools combined with specific analysis tasks and that the primary challenge as we attempt to improve will be determining the proper mix of manual and automated solutions.  After all a fool with a tool is still a fool!

A simple example will illustrate this point.  At one point in an organization’s modernization program, someone realized that much of their data was poorly stored in the clear text/comment fields of their old system. It thought that a manual approach would be required to clean and restructure the data to prepare it for use in the new system. As simple set of calculations indicated that, the time required to implement this manual approach to data quality engineering for approximately 2 million SKUs would run literally into person centuries.

Instead, a combination of automated processing was able to reduce the "problem space" from a 100 percent manual approach to a much smaller task requiring manual attention to less than 7.5percent of the original inventory. More importantly, we were able to demonstrate that we could objectively identify the point of diminishing returns – where more work on the automated approach did not produce a greater time/effort savings.  This kind of synergistic approach is common to most data quality engineering challenges.

Given the above, it remains clear that the best approach to resolving some of today's data quality challenges is to form specialized data quality teams dedicated to resolving challenges wherever and whenever they occur. Only in this manner can an organization effectively concentrate its strengths into a process that can be matured from heroic, to repeatable, to documented, to managed, and finally to improvable. Failure to do so will dilute the intellectual strength of data quality engineers with respect to their subject matter knowledge, their tools expertise, and their ability to select and apply appropriate automated solutions to specific challenges.

Until we take a unified, socio-technical approach, organizations will continue to make what they think are good decisions using bad data. I will relate one last example of a health care organization whose director indicated to the medical staff that the facility would be investing an increased amount in “knee surgery” based on an analysis of the admission data.  Unfortunately, it had to be pointed out that knee surgery was the default admission code and that the data did not reflect the true number of this type of surgery. Bad data analyzed by a well meaning administrator almost resulted in a poor decision for the organization.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.