Friday, April 5, 2013
Keeping score ...
[slideshare id=21759685&doc=dataqualitydashboards-130523061418-phpapp02]
Tuesday, March 26, 2013
No Silver Bullet for Data Goverenance
Big take away for me this week at #GartnerMDM is that technology is a tool & YOU are the enabler to MDM & Data Gov.
— DataQualityChronicle (@dqchronicle) March 21, 2013
As I mentioned in the tweet, people are the key enabler in MDM and Data Governance initiatives. I met a lot of people at the Gartner MDM Summit this past week who seemed to be searching for technology that would be the enabler, or silver bullet.
In my experience, this is an incorrect way of looking at technology. Technology is just a tool. People can use the tool to enable progress in the initiative and, in that way, technology can be enabling. But people need to be the driver and accept that responsibility.
Here are a few ways I suggest going about doing so:
- LISTEN to the business about what's going wrong and correlate that to data so you can take action
- CORRELATE the business anecdotes to understand a core list of root causes
- DEFINE the effects of the root causes
- MONETIZE the effects
- BUILD a program around correcting the root causes
This short, simple list begins with a very important point that I don't hear a lot of IT colleagues talking about. It is essential to listen to the business to learn where the pain is and what kind of issues it is causing. Don't just listen to the complaints. Drive these anecdotes to hard facts, rooted in data.
You can buy all the technology you want, but unless you perform proper business analysis, you will struggle to gain meaningful insight and direction in your data governance program.
Wednesday, October 10, 2012
The Big Data Hammer?
I've got my hammer, where is your nail?
With all the tech-media hype around big data I am reminded of the proverb ...
if all you have is a hammer, everything looks like a nail
Basically the proverb is an attempt to sum up the human condition of being single-minded and biased toward your own expertise. This is where I fear we are headed with the "big data" explosion.
Insert bold assertion here
I am of the opinion that ...
the software and consulting industry is viewing "big data" as the new technology hammer and all they can see are nails
Don't get me wrong, I see the valid use cases for a new way of doing things. I am also as excited about this as everyone else seems to be. However, I think there are some aggressive claims out there that scare me on behalf of those who are charged with managing data.
Here is an example ...
[gallery id="2483"]
[ref] Big data: The next frontier for innovation, competition, and productivity, May 2011 | by James Manyika, Michael Chui, Brad Brown, Jacques Bughin, Richard Dobbs, Charles Roxburgh, Angela Hung Byers, http://www.mckinsey.com/insights/mgi/research/technology_and_innovation/big_data_the_next_frontier_for_innovation [/ref]
Dwarfed, perhaps intentionally, by the larger, more attention grabbing graphics was the keyword could in every one of those claims. And at the bottom of the graphic was another favorite of mine ...
There seems just a few issues with capitalizing on "big data". Things like talent, which is in extremely low supply, and the legality of using inaccurate data with negative consequences (Inaccurate data? what? <--sarcasm)
Don't just take my random interpretation. Here is a great deck on when to and, more importantly, when not to use a NoSQL (aka "big data") solution:
[slideshare id=10487652&doc=gotoslides-mailable-111206135716-phpapp01]
There are some great slides towards the end of the presentation (slide 31+) that cover when and when not to use Riak, a NoSQL database. I think this quote sums up my main point ...
Don't use Riak ... if you don't have problem right now
This is a great point, so let's repeat it!
If you don't have a problem right now with your current SQL based solution, don't throw "big data" technologies at it. [ref] Perhaps a more important and subtle point in that statement is to know what problem you need to solve before you jump into the "big data" deep end. [/ref]
In other words
Make sure you need to hit a nail before you swing a hammer!
Sound off!
[polldaddy poll=6598510]
Saturday, September 29, 2012
Many Companies Unprepared for Big Data Boom
Big data has rushed on to the information technology scene over the past two years and it seems nearly everyone is looking to cash in on the promises of newly generated revenue streams teeming with profit. However, as more and more organizations make the decision to take the big data plunge, many of them are coming to the realization that they are unprepared.
Wrestling with Big Data is going to be the single biggest IT challenge businesses face over the next two years. By the end of that period they will either have it right or they will be very seriously adrift from their own business and the threats and opportunities posed by Big Data.
- Luigi Freguia, senior vice president, Systems, Oracle EMEA
Big Data, Big Demand
Like any technology solution, the return on implementing a big data solution does not come without some investment. Since big data solutions are predicated on processing very large data sets of structured and unstructured data which is very rapidly changing, one of the most critical components to the investment is hardware, lots of hardware. To give some perspective on how much data can be involved in a big data solution, Wal-Mart handles more than 1 million customer transactions every hour which they subsequently import into databases estimated at more than 2.5 petabytes (the equivalent of 167 times the books in America’s Library of Congress). This would not be possible without a significant investment in hardware.
Failure to plan is planning to fail
As I stated earlier in the article, big data has rushed on to the information technology scene and in that rush many organizations have not been able to sufficiently plan for this surge in new hardware demands. As a consequence, there has been a trend in leveraging external data centers to accommodate the demand. In a recent press release by Oracle[ref]Walker, S (2012 January 11). Oracle Next Generation Data Centre Index Raises Questions as to the Preparedness of Businesses for the ‘Big Data’ Boom. Oracle.com Retrieved January 27 2012 from https://emeapressoffice.oracle.com/Press-Releases/Oracle-Next-Generation-Data-Centre-Index-Raises-Questions-as-to-the-Preparedness-of-Businesses-for-the-Big-Data-Boom-26db.aspx[/ref] [references /], a study of approximately 1,000 managers in large organizations revealed that there was a 16% increase in organizations using an external data center. 92% of those who responded felt they would need a new data center within the next two years.
Conclusion
Big data has come on fast and furious with demands that far exceed the resources of most companies. Without an implementation plan which includes a significant investment in hardware, many organizations will find themselves unprepared for the big data boom.
Monday, April 9, 2012
Data Discovery: a path to better ETL development
Why data discovery leads to better ETL design
Let’s start with why I feel this way. Before I’d even heard of data quality I was doing it on a daily basis. You see I spent several years as an ETL developer on many data warehousing implementation projects.
Typically after a couple of briefing meetings, I’d start developing ETL mappings. Like any development effort that was followed by some unit testing where I would discover that although my ETL was written to specifications, the load didn’t “look right”.
After some digging I usually found the culprit was the fact that the source data did not match the expectations I had going into the development effort and it was time to, at the very least, add some transforms to the mapping to accommodate for the discrepancies. In effect, I was performing two critical functions left out of the original development plan, data profiling and enhancement. I feel strongly that had these two processes not been left out, I would have had a more complete and accurate ETL development experience from the get-go.
Unfortunately this was not an isolated event and, in fact, happened on almost every ETL project. First hand experience is why I feel so strongly that data discovery leads to a better development process and, ultimately, outcome.
Supporting Research for Data Discovery in ETL Design
In a fairly recent polling exercise the Passioned Group, an analyst and consultancy company, based in The Netherlands, specializing in Business Intelligence, Data Integration and ETL tools, conducted a polling of 2,000 participants where they ask what they thought were the most important requirements when choosing an ETL tool. The results demonstrated just how important data discovery is to ETL developers.
Way back in 2001 William Laurent of Information Management wrote a piece entitled, Best Practices for Data Warehouse Database Developers. The number one best practice was make sure you are provided with a usable data dictionary before starting heavy-duty development. Data discovery can help build that data dictionary without relying on assumptions and assertions made by business analysts and database administrators.
In defining what ETL is the Passioned Group mentions data profiling by explaining how it can help build a system that
that is robust and has a clear structure.
The Data Warehouse Information organization , a site “Powered by "DWH Professionals", "DWH Enthusiasts" and People alike” graphically depicts data profiling in their recommended ETL design process.
Here is an important statement they make about the benefits of data profiling during ETL design.
Data Profiling is a process that familiarizes you with the data you will be loading into the warehouse/mart
So how do I use data discovery to achieve a better ETL design?
As I mentioned in my previous post, I recommend starting with the following question:
What are the critical data domains we are looking to integrate into the target?
The reason I start with this seemingly basic question is so that you can build true discovery processes into the ETL design. True discovery finds data unbeknownst to the data consumer that also needs to be included in the target. To me, this is one of the most value added services that the ETL team can provide to the data consumers. Here is an example, taken from my previous experience, that demonstrates what I mean.
I had a marketing client that was looking to build a repository from which they could perform campaign management and analytics. They had done a fair mount of quality due diligence and identified what they felt were the required sources.
When I asked my generic question there was a fair amount of dissent in the room and some even pointed to the source to target matrix (STTM) as my source of information. However, I pressed on and discovered that some of the more executive users of the analytics were interested in performing analysis on customers were were marketed to but the address of record, for which the source systems was included in the STTM, was not deliverable (or was returned by the USPS).
As it turns out, this information was not stored in a source system but rather kept in a spreadsheet (of course) by one of the marketing administrators. Of course knowing this allowed me to incorporate the spreadsheet in the ETL sources but it also help us build in another process which discovered and profiled address data in critical business applications which were then included in an enrichment process so that undeliverable addresses could be updated with the proper addresses (where applicable).
Data discovery is a simple process once you know where to point the discovery tool. This focus is obtained by asking the general but effective question I mentioned above. Data domains, like address, help you ask more intelligent and specific questions like …
what critical applications store, collect or consume address data?
Once this is uncovered, data discovery works much the same way that data profiling works. You define the source, build a connection, define and execute the profile jobs and decipher the results.
Data Discovery for ETL Tips
Here are a few tricks I use when performing data discovery for an ETL design proof of concept.
- Profile early and often
- Translate data profiles into a metadata dictionary
- Identify data anomalies
- Never develop an ETL map from a specification, do it based on profile results
- Communicate where metadata and data distributions do not match the businesses expectations and look for the root cause
I know this list seems basic, but you’d be surprised how often it does not happen and how much rework and cost is incurred as a result.
Your thoughts?
Monday, April 2, 2012
The many uses of data discovery
the discovery of relationships between data elements, regardless of where the data is stored.
If you expand your mind beyond the conventional relational database meaning of relationships, I agree with this definition. Relationships in this context, or rather the context I chose to apply, means much more than a primary – foreign key relationship.
In this context relationships is defined as commonality. This commonality can be of a data type, value pattern, or business use. If you can profile data and understand the relationships you can set yourself up for more efficient data management practices in the areas of ETL, MDM, and application lifecycle (or application retirement). Let’s take each of these and examine how a data discovery can increase the quality of the effort.
ETL and Data Discovery
Classical ETL takes data from a source and loads it to a target. If you perform data discovery profiling on the sources before you build the ETL mapping you can achieve the following:
- a more accurate picture of the required data type of the attribute
- By examining the profile you can determine if the assigned data type is most appropriate for the data element
- a more accurate specification for the type of transform required
- If the data and metadata are not 100% coordinated you can build transforms to accommodate for this
- identification of data anomalies and outliers which require further investigation for possible remediation prior to the migration of data
- this leads to a more robust error handling and exception handling process
- the identification of data, previously unknown, that meets the business requirements and needs to be migrated
- discovery can lead to uncovering data that was previously undefined or unobtainable for data migrations
As discover tools mature, it may also be possible to generate ETL mappings directly from the tool. If the target is more richly defined in the discovery tool and the sources are more accurately identified, it makes sense to me that a discovery tool can build a better ETL mapping.
This will require a tight coupling between the discovery and ETL tool, however, there are vendors in the market with this type of coupling available to them.
MDM and Data Discovery
In the same way that data discovery can aid ETL, so too can it aid the efforts of an MDM implementation. Since MDM implementations are so dependent on ETL, the same leverage is available and can lead to a better MDM hub definition and ETL specification.
Here too can a feature to generate a data mapping be particularly useful. With so much configuration required for match and merge rules, cutting some development form the scope of the effort would only add benefit.
Another particularly interesting feature would be the ability to generate candidate schemas for the MDM hub based on the data and metadata obtained in the profiles of the sources.
Application Lifecycle and Data Discovery
Finally, during a data discovery investigation it is possible to segment data by data ranges derived from last create / update dates. This can be leveraged to perform application and/or data lifecycle management which would basically archive data past a certain date line or retire an application which has not be accessed in a predetermined, business driven date.
Here is another use for dynamically generated data mappings which would migrate the retired data to a target or archive destination.
Discovery is only the first step
As you can see from this quick summary, there are many uses for data discovery and as the tools mature there are many more things that can be done to leverage a discovery effort.
Your thoughts?
Monday, December 26, 2011
Big Data: What are some of the slick tools available to wrangle "Big Data" datasets?
What are some of the slick tools available to wrangle "Big Data" datasets?
Friday, December 23, 2011
Before you decide you want a "Big Data" solution
[slideshare id=10487652&doc=gotoslides-mailable-111206135716-phpapp01]
Antony Falco, COO and Co-Founder of Basho, put this presentation together. In it Antony talks frankly about the shift from traditional relational data storage to the NoSQL solutions such as Riak.
Here are my favorite quotes:
Would you trade ...
- Your current familiar consistency for ... a somewhat alien, but perfectly safe, consistency model and better availability
- Storage space for low latency
- late night heroics for high availability
- 35 years of RDMS success for technology used by a few companies of which you may have heard
If anyone makes these trafe-offs seem easy ... they are lying!
Any new app must use this newfangled NoSQL
Remember at a small-scale everything works ... at a large-scale, things start to break
I hope you enjoy the presentation as much as I did!
Sunday, December 18, 2011
Big Data: Big Topic, Big Confusion
Big Data Confusion
Big Data has been marketed as the solution to increase profits and aid in the discovery of all kinds of new associations from fraud detection to patient health care. Big data seems to be linked to websites like Amazon, Google and Twitter. Big Data also seems to be linked to solutions like Hadoop and BigTable. Big Data has a very different approach to data storage which includes the lack of a set schema. Big Data not only encompasses traditional data, but also includes the storage of unstructured data like documents and web page content. Because of all this the message of big data seems to get diluted, possibly even lost.
Questions seem to be more abundant than answers. Questions like :
- Is Big Data only suitable for web based solutions?
- Is Big Data only for social media?
- How can a "traditional business model" leverage Big Data?
- Who can I engage for the purchase of my Big Data software?
- How do I leverage unstructured data?
Big Data Breakdown
Big data is such a complex topic that I feel the need to start with some basic concepts and break down the details in several posts. First, I will attempt to define the big data concept, then move on to the basic principles. Last I will finish with how all these concepts drive the implementation. Keep in mind that each of these topics could cover several pages of content. Who wants to read that much? I know I don't, so I'll try and strip down these components to the bare necessity. Clarity with brevity!
Big Data 101
Big data is a term applied to data sets whose size is beyond the ability of commonly used software tools to capture, manage, and process the data within a tolerable elapsed time. These data growth challenges as being three-dimensional, i.e. increasing volume (amount of data), velocity (speed of data in/out), and variety (range of data types, sources).
Big Data 102
- Tolerance
- Availability
- Consistency
Tolerance refers to the ability to add more resources ,i.e., hardware, to the solution. Data growth is well documented and is a condition that needs to be incorporated into any big data solution.
Availability refers to the fact that Big Data solutions need to be able to guarantee that each request receives a response. Big Data solutions are often applied to web-based applications which tend to have millions of concurrent users.
Consistency refers to concept that each of the data stores in a big data solution contain data in the same state. With millions of concurrent users and the demand for availability, a consistent data state is probably the most difficult principle to ensure.
Big Data 103
These three characteristics drive the way that a Big Data solution is architected. In order to facilitate the volume and velocity, Big Data solutions need to be distributed over many environments with independent but identifical hardware and software resources. This architecture is commonly referred to as massive parallel processing, or MPP. Distributed architectures spread data over many different data stores according to which environment is available.
This leads to one of Big Data's most contraversial consequences; eventual consistency.
Eventual consistency boils down to the fact that, with a Big Data solution, there are points in time when the data in the environments are not consistent. Environments are eventually synchronized when resources can be made available to do so.
Eventual consistency will be a topic that I plan on elaborating on in future posts as I feel it is the tie between big data and data quality.
Summary
Big Data is not about a website, vendor or social media alone. It is a unique way to store and continuously deliver massive amounts of diverse data. It is predicated on three principles, often referred to as CAP, which ensure that the solution is scalable and available, but tends to sacrifice data consistency. In order to do this, Big Data solutions require a distributed set of resources. While this may sound like a watered down explaination, there are many details that I will elaborate on in future posts.
Friday, August 26, 2011
ABC and DQ: Codependent Initiatives?
Activity Based Cost and Data Quality: Codependent Initiatives?
Summary
Activity Based Costing, or ABC, is an exercise where costs are assigned to business activities required to support critical business operations. While it is often used in support of a business process redesign (BPR) effort, it can also serve an important role in data quality (DQ) initiatives.
In order to conduct a data quality initiative, a significant investment is required. There are costs to purchase hardware, software, support and implementation resources. Considering DQ efforts are usually associated with "fixing" a previous investment, these costs are not generally accepted as capital investments. They are usually viewed as negative consequence to a failure to comprehensively implement a system.
One way to mitigate this perception is to demonstrate, in monetary terms, the return likely to be realized. Demonstrating this return is most affective by linking cost avoidance associated with increased business operations. In order to do this, you need costs associated with these operations.
As a result, DQ becomes dependent on ABC in order to justify the expense.
The proof is in the pudding. And by pudding, I mean operation
So let's talk about ABC efforts and how they can be used to help justify why an organization needs to invest in data quality. Rarely, does an organization know how long it takes someone to perform everyday business operations. After all, this time is often viewed as a necessary expense of doing business and anytime invested is required so why analyze it? Given this approach, it is not surprising that very little is known about how long it takes to perform a business operation when there are exceptions to the norm.
In other words, when there are data related discrepancies very little is tracked in terms of impact and cost. However, this is a legitimate and every day reality. Often reports do not reflect the same aggregation and someone spends countless hours tracking down the root cause for the discrepancy. This time is money. It is also lost opportunity, which results in further cost to the organization. Boiling this type of event down to a cost of resolution can form the basis for justifying investments in prevention.
For example, if there is a report that provides the status of product inventory and another that provides a summary of sales these two reports should decrease and increase in direct proportion. Inventory should decrement at the same rate as sales increments. The chart below illustrates this relationship.
- Inventory vs. Sales Report Relationship with High Quality Data
This doesn't take long to look at before you start realizing there is a problem, especially for someone who is responsible for generating this report and is familiar with much more normal looking analysis. Logically, this individual will start down the path of determining the root cause. At this point, the meter starts running. However, not only one meter but two are racing toward an unexpected cost. One of the meters tracks the time, and hence money, spent fixing the issue. The other meter tracks the time and money not spent performing duties that would have been performed had the issue not been there in the first place.
If you boil down this individuals compensation down to dollars per minute and track the time taken to resolve this issue and time not spent producing in other areas, you can start calculating the cost of poor quality data. In all liklihood this individual does not resolve the issue alone, so you can start adding in additional dollars a minute for supporting members. Not to mention, sales people don't have accurate inventory numbers and this impacts their ability to maximize their activities! Before long, you start to get a clear picture that poor quality data is costing you, in a recurring nature, a lot more than the cost of fixing it.
Although disturbing, this is the key to ending the vicious cycle of waste and becoming a more streamlined organization.
Conclusion
It is a difficult decision to spend money on fixing broken operations and systems, however, it is an easy decision to spend money to end waste and increase productivity. It just a matter of perspective. An experience data quality professional will help you see this effort in the right light and even help you put real numbers behind it.
If all this sounds familiar, maybe it is time to find that data quality professional and start saving some money!
Monday, June 13, 2011
The role of data quality in ETL design: DQETL
Introduction
Data integration is nothing new. Since the concept of data warehousing, data integration has been a major initiative for most large organizations. On the most common obstacles of integrating data into a warehouse has been the fact that assumptions about the state of the source data have been either false or flawed at best. One of the reasons for this is that very little investigation, or data profiling, is performed on the source data prior to design and execution of the data extraction routines.
With all there is to plan for and deliver on data warehousing projects, this oversight is understandable. However, I believe, it is time for data quality to assume the role of reporting on and remdiating the current state of all source data to be migrated into a warehouse.
Turning assumptions into facts
If source data was profiled what was once assumptions about data can be transformed into facts about the state of the data . Data profiling consists of scanning data and typcially delivers measures such as the frequency of nulls, the extent to which data is unique, and ranges of actual values within each fields included. With data quality tools such as Informatica's Data Quality v9, Global ID's data profiler or Talend's data profiler, these basic reports can be compiled with a few clicks on the mouse. Furthermore, this products offer portals where dashboards detailing the current state of the data can be delivered to both a technical and business oriented audience.
Data Profiling 101
As I mentioned, creating a data profile can be done with a few mouse clicks. Typically the steps are as follows:
- Define a connection to the data source
- Define the data source
- Define which fields are to be included in the profile
- Define any business related rules to be included in the profile
- Schedule the profile for execution
Connections
Defining the connection to the data source usually involves a few simple steps. Connections are typcially either to a database or to a file. While connecting to a file includes parameters such as delimiters, field names, and data types and lengths, connecting to a database usually involves location parameters like host and authetication credentials like username and password.
Whether you connect directly to the database or use a flat file extract typically depends on circumstances like resources for a test environment and ability to procure the required credentials. Either way, the lion's share of the work is setting up the connection.
Candidates
The next logical step in creating a data profile is to define what fields to include. Within the context of validating ETL for data integration, this step would heavily depend on those fields nicluded inthe migration, and even more by those fields requiring transformation.
Bercause not all data is migrated, data profiling is best of limited to the tables that are required. Primary and foreign keys are almost always included to ensure uniqueness as well as timestamp fields to ensure completeness and date format conformity.
Rules
When we talk about data profiling most of the metrics are technical (percedntage of uniqueness, percentage of nulls, etc ...), however, once fo the most beneficial practices is to include rules in the data profile that are based on business rules. Some examples of these rules are to validate that certain chronological events are in order (i.e. ship date does not proceed order date) and logical assumptions (i.e. that individuals who are indicated as male do not have postive pregnancy outcomes) are indeed valid.
Constructing business rules often involve participation of a business domain subject matter expert, however they can also be formed from the conceptual deisng of the ETL. By reverse engineering the transformation logic, it is possible to derive, at least, one rule that needs to be tested. Afterall, transformation logic would be negatively affected by things such as nulls, data type nonconformity and values outside the expceted range.
The upside of yet another step
Because of its complex nature and multi-step requirements, adding another step to data integration and migration is rarely a welcomed effort. However, once viewed within the context of reducing ETL redesign and increasing transfer success rates, it is frequently, albeit begrudgingly, accepted.
Including data profiling in the data maigration suite of perations, indeed, can deliver these desired outcomes. When data is profiled prior to ETL design and execution, data states that would otherwise cause ETL loads to fail can be identified and remedied. An example can be found in the all important date related fields. A simple data profile can detect date formats that are not supported by the target repository. Dates in data warehousing are crucial to track transaction lineage and if not configured correctly can be the undoing of an ETL design.
Summary
While I have just touched the surface of the role of data profiling in data integration this is, none the less, an important concept to adopt. For increases in successful loads and decreases in ETL test and troubleshooting will save time and resources and paint a more positive image of the data integration team and their capabilities.
Monday, June 6, 2011
Big Data ... Little Data Quality
Is Big Data better Data Quality?
Big Data is everywhere. Chances are you've used a big data solution today. However, are big data solutions delivering big data quality?
High Availability versus High Data Quality
Typically, Big Data solutions are designed to ensure high availability. High availability is based on the concept that it is more important to collect and store data transactions than it is to determine the uniqueness or accuracy of the transaction. Some common examples of big data / high availability solutions are Twitter and Facebook.
It is possible to configure a big data solution to validate uniqueness and accuracy. I want to make sure I state that clearly. However, in order to do so you need to sacrifice some of the aspects of high availability to do so. So, in some regard, big data and data quality are at odds.
This is because one of the fundamental aspects of high availability is to write transactions to whichever node is available. In this model, consistency of transactional data is sacrificed in the name of data capture. Most often, consistency is eventually configured on data inquiries, or on data reads as opposed to data writes.
In other words, at some given point in time you do not have consistency in a big data dataset. Even more troubling is the fact that most transactional conflicts are resolved based on timestamps. Which is to say that the most recently updated transaction is commonly regarded as the most accurate. This approach is, obviously, an issue that requires further examination.
Room for improvement
As we examine big data solutions and learn more about implementing them, it is important to design more robust conflict resolution approaches that ensure that big data includes big data quality.
More on that to come ...
Friday, May 20, 2011
The Seven Habits of Highly Effective Data Quality
7 Habits of Highly Effective Data Quality
I've been reading Stephen Covey's The 7 Habits of Highly Effective People and I couldn't help but notice the parallels between effective people and effective data management. In the book Covey discloses that there are principles, centered on self-discipline, that lead to success and fulfillment. Sounds great, right?
The seven habits include some ear-cringing buzz words, but let's take a look at them and their data quality doppelgänger.
Be Proactive
For years data quality has been a discipline striving to transform itself from reactive to proactive. In fact, the ROI in data quality programs centers on being more proactive to avoid regulatory issues and costs and improving decision making. It's an understatement to say that data quality programs need to be focused on taking the initiative and become proactive programs of change.
Proactive data quality means identifying and remediating data quality issues before they become proliferated throughout the enterprise. Simply put, proactive data quality is about having identification and remediation processes at data entry points and addressing issues at the source.
Begin with the end in mind
Beginning with the end in mind brings a smile to my face. This was practically the title of one of my first posts for this blog. Without knowing where you need to end, your route to that end will almost inevitably be scattered and twisted. For it is only by setting a clear destination that a clear path can be developed. Often, in the world of data quality, setting a destination focuses on developing metrics and targets that will bring about positive change in the organization.
Put first things first
Putting first things first is about setting priorities and building a course of action(s) that will address the prioritized list of objectives. In others words, don't focus on everything all at once but rather break down large tasks into smaller more achievable parts. This is often useful when developing and implementing data quality programs because there are so many moving parts that need to be put in place simultaneously.
Think Win-Win
Win-wins in the data management / data quality arena are all about implementing rules that help multiple business units improve their data and its use. There are some easy domains where one data quality service equates to a win-win.
Address validation is a prime example of the win-win scenario. Every business unit benefits from more accurate customer addresses. Implementing address validation processes can be orchestrated in such a way that the process can accept different address sources and implement the same validation routines. Not only is this a win-win, it also cost effective and generates a high rate of return on investment.
Seek First to Understand, Then to Be Understood
This one is pretty straight forward. Data quality / data management is all about solving problems and building effective change. You can’t be affective at solving a problem without first knowing what it is. A more subtle point I’d like to make here is that all too often there is a tendency in the technology field to explain the intricacies of the solution. Frankly, business people don’t care how you solve the issue just that you do solve it accurately. Only understanding issue ensures that you can do this.
Synergize
Cringe! Worst buzzword ever? Maybe. In essence synergy means bringing together a whole that is greater than a sum of its parts. As described in the win-win section, synergies in data management / data quality are largely derived from building a solution that works for multiple business units in such a way that they produce a benefit greater than if the solution was only built for one unit.
That said, building a solution that “chains” several beneficial processes together like address validation and duplicate reduction can also be thought of as a way of bringing together a whole greater than the sum of its parts.
Sharpen the Saw
My personal favorite! Sharpening the saw has to do with the continuous process of developing skills. In part due to the wide range of data quality modules, there is always a need to sharpen the saw. For example, I am currently working on expanding my ability to produce more accurate matching techniques so I can be sure that I identify true duplicates and produce the minimal amount of false positives. In addition, I am always searching for more knowledge on address validation techniques.
Sharpening the saw with regard to data quality processes is a way to revisit the existing solution and make it better. This is an essential practice due to the growing number and varied nature of data sources continuously added to the enterprise landscape.
Conclusion
Effective people and effective projects and strategies can learn a lot from Covey’s 7 habits research. I encourage those of you reading this post to try and implement these habits not only in yourself but also in your projects!
What data quality is (and what it is not)
Like the radar system pictured above, data quality is a sentinel; a detection system put in place to warn of threats to valuable assets. ...
-
Recently I had coffee with Dr. John Talburt of the University of Arkansas at Little Rock's Information Quality program . During the conv...
-
[caption id="" align="alignleft" width="240"] Image by Marius B via Flickr[/caption] Is Big Data better Data ...