Monday, April 1, 2013
Data Management Poll Results
[gallery type="slideshow" ids="2761,2759,2760,2758,2757" orderby="rand"]
Monday, June 6, 2011
Big Data ... Little Data Quality
Is Big Data better Data Quality?
Big Data is everywhere. Chances are you've used a big data solution today. However, are big data solutions delivering big data quality?
High Availability versus High Data Quality
Typically, Big Data solutions are designed to ensure high availability. High availability is based on the concept that it is more important to collect and store data transactions than it is to determine the uniqueness or accuracy of the transaction. Some common examples of big data / high availability solutions are Twitter and Facebook.
It is possible to configure a big data solution to validate uniqueness and accuracy. I want to make sure I state that clearly. However, in order to do so you need to sacrifice some of the aspects of high availability to do so. So, in some regard, big data and data quality are at odds.
This is because one of the fundamental aspects of high availability is to write transactions to whichever node is available. In this model, consistency of transactional data is sacrificed in the name of data capture. Most often, consistency is eventually configured on data inquiries, or on data reads as opposed to data writes.
In other words, at some given point in time you do not have consistency in a big data dataset. Even more troubling is the fact that most transactional conflicts are resolved based on timestamps. Which is to say that the most recently updated transaction is commonly regarded as the most accurate. This approach is, obviously, an issue that requires further examination.
Room for improvement
As we examine big data solutions and learn more about implementing them, it is important to design more robust conflict resolution approaches that ensure that big data includes big data quality.
More on that to come ...
Thursday, March 24, 2011
Data Quality Polls: Troubled domains and what to fix
As expected, customer data quality remains at the top of list with regard to having the most issues. Ironically, this domain has been at the forefront of the data quality industry since its inception.
One reason for the proliferation of concerns about customer data quality could be its direct link to revenue generation.
Whatever the reason, this poll seems to indicate that services built around the improvement of customer data quality will be well founded.
[caption id="attachment_1049" align="aligncenter" width="630" caption="What would you improve about your data?"]
Once again there are no surprises when looking at what data improvements are desired. Data owners seem to be interested in a centralized, synchronized, single view of their data, most notably customer.
The good news that can be gathered from these polls is that as an industry, data quality is focused on the right data and the right functionality. Most data quality solutions are built around the various aspects of customer data quality and ways to improve it so there is a master managed, single version of a customer. The bad news is we've had that focus for quite some time and data owners are still concerned.
In my opinion, this is due to the nature of customer data. Customer data is at the core of every business. It is constantly changing both in definition and scope, it is continuously used in new and complex ways, and it is the most valuable asset that an organization manages.
One thing not openly reflected in these polls is that it is likely that the same issues and concerns that are present in the customer domain are also present in the employee and contact domains. However, they tend not to "bubble up" to the top of list due to lack of linkage to revenue and profit.
I'd encourage comments and feedback on this post. If we all weigh in on topics like this, we can all learn something valuable. Please let me know your thoughts on the poll results, my interpretation of the results and opinions.
Wednesday, December 22, 2010
Putting your best foot forward: Data Quality Best Practices
Data Quality Program Best Practices
Implementing an enterprise data quality program is a challenging endeavor that can seem overwhelming at times. It requires coordination and cooperation across the technology and business domains along with a clear understanding of the desired outcome. A data quality program is fundamental to numerous other enterprise information management intiatives, not the least of which are master data management and data governance. In fact, you'll recognize some of the same best practices from those disciplines.
1. Establish Sponsorship from the Business
Gaining business level sponsorship for the data quality program is essential to its success for many reasons. Not the least of which is the fact that poor data quality is a business problem which negatively impacts business processes. Sponsorship from the business provides a means for the communication of these problems and impacts. Business level sponsorship should be built upon data quality ROI and provide the direction for what data resides in the scope of the data quality program.
2. Establish Data Stewards
Data stewards are business level resources that represent individual data domains and provide relevant business rules to form remediation steps. They also develop business relevant metrics and targets that enable the measurement of quality and trend reporting that establishes the rate of return on existing remediation techniques.
3. Establish Data Quality Practice Leader
Data quality is not a one-and-done project. It is a cycle of activities that need to be continuously carried out over time. Enterprise data tends to be an evolving and growing asset. Assigning a leader, or set of leaders, to a data quality program ensures the data quality cycle of activities maintains consistency as the data landscape undergoes this evolutionary growth.
4. Establish Data Correction Service Level Agreements
Defining service level agreements with the business data stewards provides a basis for operational prioritization within the data quality program. As new data defects are discovered it will be critical to determine which to schedule for implementation. Service level agreements provide direction from each business unit and enable appropriate change management scheduling.
5. Tightly Couple Data Quality Processes with Data Management Practices
Common data management practices such as data migration and data archive scheduling need to be taken into account when determining when and what data to assess and remediate. Profiling and assessing data which is scheduled for archival would be an egregious misappropriation of resources. By aligning the data quality program with these types of data management activities this type of mistake can be avoided.
6. Approach Data Quality Proactively
A proactive to data quality increases data consumer/business confidence, reduces costs associated with unplanned data correction activities and fines associated with failure to meet regulatory compliance. A proactive approach also establishes, with the business, a level of domain expertise that fosters the necessary buy-in. Fundamental to this concept are data quality assessments and data quality trend reporting (score-carding).
7. Choose Automated Tools over Manual Methods
With the maturation of the data quality vendor market, it is now possible to implement enterprise capable data quality software that offer a full range of features to management data defect identification, remediation and reporting. Automated tools are more comprehensive, consistent, portable, and include built-in modules, such as address validation services, which reduce code development.
8. Establish Data Metrics with Targets
Metrics, and their associated targets, are the cornerstones to developing an assessment and remediation process that fosters affective change. The definition of metrics and their targets needs to be centered on data elements that are essential to the core set of business processes. Targets should be divided into three groupings which are reflective of ability to support these processes. At a base level these groupings should be “does not support business function”, “minimally supports business function”, and “supports business function at a high level”.
9. Address Data Quality Issues at the Source
With numerous applications consuming and delivering data across the enterprise, it can be a daunting task to decide where to start correcting quality deviations. In an effort to reduce this complexity, it is a best practice to institute data quality activities where the data originates. The origin point of data is commonly referred to as the system of record.
This practice not only provides an answer of where to begin implementing data quality practices, it also proactively addresses the issue of defect proliferation. This ensures that these activities are not duplicated numerous times and are implemented in a consistent manner. As a consequence, data needs to be measured for quality upon creation and/or migration into the data landscape.
10. Focus on Mission Critical Data
Focusing data remediation efforts on mission critical data is the control that ensures a return on the investment of the program. Identifying the data that support core business functions requires careful examination of the process and participation from the business data stewards. Often times this process also requires prioritization of critical elements in order to schedule remediation efforts. The identification of this data is vital to the success of the data quality program.
Summary
While there are many more best practices in the data quality domain, these ten form a solid foundation for the implementation of a data quality program. This practices, as you may notice, are more focused on establishing a data quality program rather than the remediation efforts within the program. In a future post, I'll examine some common remediation techniques which are universal to data quality programs.
Saturday, September 11, 2010
Master Data Management: Address Validation Series: Address Validation
Why?
There are plenty of aspects of address validation to write about.
Validating addresses can be done with many different tools, each with their own specific details on how to do it. There are various ways to validate address within each tool to produce different outcomes. And there are various ways to manage and integrate this data back into enterprise, operational systems.
While all of this content is very helpful and important to convey, it is my belief that it is essential to understand, write about and discuss why address validation is important to an organization and their master data management (MDM) efforts.
Here are a few reasons why address validation is so important to any organization (feel free to comment and suggest others!):
- Without valid address information, return mail can impact an organizations bulk mail status and lead to increased mailing costs
- Without valid address information, billing operations generate negatively impacting cycles of billing collections and corrections which has a negative impact on revenue assurance
- Without valid address information, marketing campaigns are not fully leveraged
- Without valid address information, marketing techniques like house-holding cannot fully realize their potential
- Without valid address information, customer care operations are impeded
- Without valid address information, customer perception and the customer experience is negatively impacted
- Without valid address information, an organization is open to federal regulatory fines resulting in the failure to honor contractual obligations
- Without valid address information, shipping operations experience failures which generate negatively impacting cycles
- Without valid address information, supply chain management efforts are comprimised leading to reduce costs and increase effectiveness
- Without valid address information, asset management efforts for business models built on locational awareness, like housewares rental providers and home security providers, are seriously comprised
I'm sure there are other reasons, but the key message to convey is that validating address information is a critical component to business operations throughout the enterprise.
How?
Now that we've covered why it is important to validate address data, let's concentrate on how. While there are many tools on the market with which to validate addresses, I use Informatica's v9 product which integrates AddressDoctor v5 to accomplish validation.
The basics
In order to validate address data with Informatica there are three required components:
- An input containing address information
- The Global Address Validation (GAV) component
- An output to write the original and validated address information
Address Inputs
As far as address inputs are concerned, no surprises here. Typcial address information such as street number and name, city or locality, state or province and postal code can be passed into the GAV module. Of these, the street number and name along with the postal code are required.
Global Address Validation
The Address Doctor address validation service performs several types of address processing which are beneficial to address validation and address matching. Among these are:
- Delivery Address validation
- Formatted Address validation
- Mailability validation
Delivery Address Validation
Delivery address validation involves verificaiton, and if necessary correction, of the street number, street name and any sublocation information. This process ensures that the address is valid and deliverable via USPS standards. This process is a must for marketing and billing operations that want to ensure mailings reach their destination.
Formatted address validation takes the various inputs provided and arranges them into the standard mailing format. By standard mailing format, I mean the way the information should be presented on a mailing envelope. This process is particularly useful to marketing operations to process "raw" address data into suitable mailing data. Refer to the illustration below for an example of formatted address validation.
When it comes to address validation perhaps the most important result is the verification of the deliverability of the address. The GAV offers a way to investigate and report on this critical aspect of an address. Below is an illustration of address validation results including address deliverability. As part of this process there are match codes provided that range from validating the address as deliverable to stating the address could not be validated or corrected. Refer to the illustration below for an example of delivery address validation.
Address Validation Mapplet
Since we are performing address validation in support of a master address management initiative, it is best practice to use a mapplet to perform the validations. A mapplet is a reusable object containing a set of transformations that you can use in multiple mappings. Any change made to the mapplet is inherited by all instances of the mapplet. In this way, mapplets are an object oriented approach to performing data quality.
The basic requirements for a mapplet are an input, a transformation and an output. In our address validation example the transform is the GAV previously mentioned. Below is an illustration of a basic address validation mapplet.
Here is a step-by-step process on how to create such a mapplet.
Step 1) Define your input information
- 1. Click on the input component on the component toolbar
- 2. Right-click on the input component and click on the properties option
- 3. In the ports section of the properties window click on the new icon
- 4. Define the port name, data type, precision and scale
Step 2) Define transform properties
2a) Add the Global Address Validation component
2b) Define transform input parameters
1. Right click on the transform and select the Properties option
2. Select the Templates option from the Properties tab on the left
3. Click on the (+) icon next to the Basic Model option
3a. Click on the (+) icon next to the Discrete option
4. Select the following input parameters from the Discrete option: StreetComplete1, StreetComplete2, PostalPhraseComplete1, LocalityComplete1, Province1
2c) Define the transform output / validation parameters
1. Click on the (+) icon next to the Address Elements option
2. Select the following attributes from the Address Elements option: SubBuildingNumber1, SubBuildingName1, StreetNumber1, StreetName1, StreetPreDescriptor1, StreetPostDescriptor1, StreetPreDirectional1, StreetPostDirectional1, DeliveryAddress1, DeliveryAddress2
3. Select the following attributes from the Last Line Elements option: LocalityName1, Postcode1, Postcode2, ProvinceAbbreviation1
4. Select the following attribute from the Status Info option: MatchCode
5. Select the following attribute from the Formatted Address Line option: FormattedAddressLine1, FormattedAddressLine2
3) Define an Output Write object
1. Select the Output Component from the Transformation Palette
2. Right click on the write object and select the Properties option
3. From the Properties tab, select the Ports option
4. Right click on the first available row and select New from the menu options
5. Enter the name of the output field
4) Connect the Validation component to the Write Output
In this step you connect the output ports from the address validation component to the write output object. This is a simple drag-and-drop step connecting the appropriate ports to each other. If you keep your names consistent it'll help keep you sane.
Conclusion
Now you've learned how to create an address validation mapplet! This is a great step toward building a consistent, repeatable address validation process which is the key to implementing a master address data management program!
Thursday, July 1, 2010
Master Data Management: Address Validation Series
Why?
Address information, in particular customer address information, is a core asset of any business. It plays a pivotal role in two fundamental business operations; revenue assurance and revenue generation.
Without valid, deliverable customer address information collecting payment for services or products is often a process that, at best, requires repetitive efforts that cost the business labor and resources (and dollars). At worst, the process fails to collect, creating an obvious issue costing time and resources (and dollars).
Without valid, deliverable customer address information marketing to existing and potential customers is not possible and will, again, cost the business labor and resources (and dollars).
What?
So what exactly needs to be validated in order to prevent the failure of revenue collection and generating events?
While it is not harmful to have the full compliment of customer address data collected, stored, and validated, there are a few pieces of address information that are essential.
Postal Code is an absolute must have in order to ensure the mailer be delivered. Postal Code is the core element that the United States Postal Service uses to route mail. Without it deliverability is unachievable.
Street Number and Street Name are also essential pieces of information to collect and validate in order to ensure mail deliverability. Logically, this information is required to know where within the postal code to deliver the mail.
It is also required, where applicable, to collect and validate additional address location information such as apartment number, suite number, building number, etc. This enables getting the mail to correct destination within a multi-dwelling residence.
In my opinion, this is the required data to validate and ensure deliverability. Most address validation services can derive accurate and valid city and state information from the postal code which can be augmented and utilized moving forward.
Who?
Who should be responsible for address validation?
As I eluded to earlier, address information is a corporate asset which plays a pivotal role to many essential business operations.
For this reason, address validation belongs in a centralized group made up of representatives from those dependent parties. In other words, address validation is the responsibility of a corporate data governance group that is aware of all the required aspects of useful address data management.
Typically there are, at a minimum, two levels to this group. On one level there are business stakeholders that manage and advocate functional business requirements involving address information. On the other level are the data processors that manage data sourcing, scrubbing, validation and integration of the address information.
Due to the technical requirements in managing such information, Information Technology should be responsible for management of the physical data stores that house the address information. However, it is crucial to note that this management is around the software and hardware resources that house the data.
It is imperative that data ownership be the responsibility of the business owners on the governance group.
Where?
Although there are various implementations of MDM, I believe address information belongs in a a centralized hub that feeds dependent systems clean and valid address data. This model ensures the delivery of consistent and valid address throughout the enterprise.
This centralized hub needs to be managed in such a way that it is independently supported ensuring failover, redundancy and archival. This eliminates the failure scenarios described earlier that interfer in revenue collection and generation.
When?
How often does address data need to be validated?
There are various factors such as an annual change of address of 17% and quarterly marketing campaigns that influence when address information should be validated. In the end, the answer to when should address data be validated depends on the lowest level of granualarity that the data is used to support business operations that either collect or generate revenue.
If marketing conducts campaigns on a quarterly basis but billing occurs monthly, than validating address data should be done on a monthly basis to support accurate and efficient billing operations.
How?
How can address validation be implemented in order to support all the benefits described?
In order to validate address information on a periodic basis, manage it across various dependent business units and integrate it into a centralize hub you need to be able to develop validation routines, business rules, a mechanism for business stakeholder review and integration routines that can be executed in a scheduled format.
Within the domain of address validation there are several varieties of output. For instance, it possible to develop an address validation process that transforms address information into the correct formatted address lines that would appear on the envelope. Another implementation could be the parsing, augmenting and obtaining validation status of address information. Yet another implementation could be to take the address information input and transform it into the valid delivery address information.
With various business units consuming address information, there will likely be various business unit specific rules to process the address information. For example, marketing operations might require the "vanity" city name be specified. Vanity city names are usually preferred by customers due to their perception and reputation. One such example of a vanity address is using Beverly Hills over the validate city name of Los Angeles. However, billing operations may not have the same requirement. In this case, and others like it, you need an address validation process that enables the building of business specific rules that can handle variability on the same data element.
In order to enable business stakeholder ownership and help business users define and validate data specific rules, you need to have a mechanism that presents data to these users. Since these business stakeholders are not typically technically inclined, this mechanism needs to be built in such as a way that minimizes technical effort and enables data review and validation.
Ultimately this address information needs to be integrated into a centralized hub and distributed to the various consuming applications. This dictates the need for enterprise capable data extraction and load features such as scheduling, monitoring and tracking.
How do you deliver on such a complex set of criteria?
It's a challenge. In fact, it's such a broad topic with many details that it is not feasible to do in one blog post. I plan on addressing (no pun intended) each of these areas in more detail over the coming weeks.
So stay tuned to The Data Quality Chronicle for more!
Tuesday, May 4, 2010
On Cloud 9!
Apologies ...
I've been in the clouds lately, in more ways than one. I've been on the road performing another data quality assessment on an island in the Pacific. This translates into the fact that I'm gaining status on multiple airlines and becoming increasingly appreciative of noise canceling technologies.
I'm also gaining an appreciation for another technology, cloud based data quality solutions! I am leveraging Informatica's latest data quality platform, IDQ v9. IDQ v9 brings to mind a favorite 80's commercial of mine where peanut butter and chocolate are combined into one tasty treat! For sure there is a little PowerCenter in your Data Quality and a little Data Quality in your PowerCenter ...
50,000 foot level view
I won't try to cover the upgrade in one post, but rather just wet your appetite on what's inside the wrapper. We can binge on the details in the coming months. For now let me just highlight what I feel are the most significant enhancements of v9.
- Quick start solution: the cloud based solution of v9 almost eliminates the previous installation requirements of IDQ 8.x
- Data Explorer and Data Quality are now one product: This cuts installation and repository management by 50% (at least)
- PowerCenter integration means ETL and DQ have tied the knot!: Now you can stop using TCL & SQL scripting and leverage PowerCenter's integration components. This includes the consolidation component, a particularly important component to master data management and customer data integration initiatives
- Inline data viewing: now you can unit test your transforms without a full run on your mappings, saving time and increasing productivity
- Inline data profiling: now you can report on data quality processing without leaving your development environment and share it with client via a URL!
As I mentioned, I'll dive deeper when I'm not in the middle seat of row 42 somewhere over the Pacific. For now my big take-aways are the time-saving features that are almost everywhere and the integration with PowerCenter takes data quality integration to the next level.
No more afternoons installing client and scripting repositories. No more SQL development to valid and analyze results. No more TCL script (I love that!). No more SQL consolidation nightmares!
Up to speed
As for a learning curve on the new look and feel? It took me a few hours, of which most were productive hours I was able to leverage into real work. Hopefully I can translate what I've learned and cut your learning curve down with posts in the coming months. For now, the drink cart is coming my way and I've got cash handy!
Check back next month when I go over how I was able to deliver more analysis in a shorter time frame and look good doing it!
Monday, March 1, 2010
IAIDQ El Festival del IDQ Bloggers: March Edition
El Festi... What?
March Submissions
Agapo Data!
Jim Harris, Blogger-in-Chief of Obsessive-Compulsive Data Quality, submits an interesting piece this month entitled Maybe you're just not that into your data?. In this piece Jim uses a clever musical parody to remind us that "data needs love too" and that data quality project are not a one-time project but rather need to be thought of as a "sustained program". Enjoy the entire blog post here.
Jim is an independent consultant, speaker, writer and blogger with over 15 years of professional services and application development experience in data quality (DQ), data integration, data warehousing (DW), business intelligence (BI), customer data integration (CDI), and master data management (MDM). He is a member of the International Association for Information and Data Quality (IAIDQ) and the Iowa Chapter of Data Management Association International (DAMA). Learn more about Jim here.
Philia Data!
Daragh O Brien submits a telling parable, from IQTrainwrecks, of the potential costs and embarrassment that can arise if critical master data is not managed correctly and if sanity checks are not built into processes. UK local authorities have been taking themselves to court and pursuing costs against themselves for failure to pay parking fines. Read the whole story here.
A busy man, Daragh is the driving force behind DoBlog. Launched in 2006, the DoBlog is the personal blog of Daragh O Brien, former IAIDQ publicity director and information quality consultant based in Ireland. In addition Daragh is also founder of Castlebridge Associates which is a specialist Information Quality, Data Protection and Data Governance consulting practice based in Ireland. The Castlebridge Mission is to help organisations manage their information assets as well as they manage their people assets.
IQ Trainwrecks was established in 2006 by the IAIDQ, IQTrainwrecks.com is a community reference site capturing case studies and examples of the impacts of poor quality information and data in the wild.
Data Mythos!
In this submission Henrik Liliendahl Sørensen explores the mythology behind why it seems most view data quality projects as a technology issue rather than a business one. Regarding one of the more commented-on blog posts in recent memory, Henrik laments that "the best moments in blogging is having a lot of sensible comments". See what all the buzz was about here.
Another busy man, Henrik is a data quality and master data management professional also responsible for creating data architecture solutions. He currently is the Practice Manager for Omikron Data Quality. Learn more about Henrik here.
Referential Treatment
Steve Sarsfield submits a post this month regarding the value of external data in the success of a data quality initiative. From availability to Geocoding and U.S. Census data, Steve explores the exciting applications for external reference data in the validation and standardization of data. Get the details straight from Steve here.
In addition to being the driving force behind the Data Governance and Data Quality Insider, Steve is also a data quality evangelist and author of the book the Data Governance Imperative. When he is not busy writing Steve works on product marketing at Talend where he focuses on product strategy and press/analysts relations.
Integrity of Enterprise Data
Ken O'Connor explores one the more socially engineered pitfalls of data quality in his blog post where he recalls recent conversations with clients and assertions from Craigslist founder Craig Newmark. Find this and related blog posts here.
When Ken is not sharing his data management experiences in his blog Ken O’ Connor Data Consultant, he is an independent IT Consultant specializing in Data: Data Migration, Data Population, Data Governance, Data Quality, Data Profiling, Master Data Management, Business Intelligence. Ken is the founder of Professional IT Personnel Ltd. Check out more about Ken here.
In Closing ...
I've hosted the Festival del IDQ Bloggers once before and it is always an exciting opportunity to network with other data quality experts, read some great blog posts and exercise my skills in the writing forum. I encourage anyone thinking about hosting to do so. It's easy, fun and doesn't take a great deal of time to complete. For more information on how you can host the Festival contact IAIDQ's Director of Publicity, Heather Richards, here.
Monday, November 30, 2009
Microsoft Dynamics CRM Duplicate Consolidation Management
For our sample records let's say we have just two contacts that are duplicates. Contact A has four service calls associated with it. Each of these service calls have relevant data that you want to retain. Contact B has three service calls associated with it and each service call has data that needs to be retained.
Upon merging Contact A with Contact B (so in this case B is the keeper record), there will be seven service calls associated with Contact B. This is accomplished through the use of four data elements in each Contact transaction. These fields are ContactId, MasterId, Merged, and Statecode.
Merged is an indicator field where 1 indicates that the transaction is indeed merged. Statecode is another indicator field indicating active and inactive transactions. In Dynamics a statecode of value of 1 is inactive and 0 is active. Yes, you read that right. Zero is active.
The "magic" of the duplicate consolidation lies in the MasterID field. For consolidated records the MasterID is equal to the unique identifier of the keeper record. So in our example if the unique identifier of Contact B was 1234, the MasterID of Contact A will be 1234. The Merged field would be 1 and Statecode of 1.
In addition to this update, there is another set of updates required in order to "reassociate" each transaction formerly associated with the subordinate record to the new master.
So continuing with our example, the service call entity would need to be updated so that the transactions with the ContactId of the non-master will need to be updated with the master ContactId.
For those transactions which are not merged the MasterID is NULL (no need to store the unique identifier twice).
Saturday, October 17, 2009
Removing duplicates in Microsoft Dynamics CRM
Before I get into the details, I want to emphasize that without customization, removing duplicates is not a batch process. In other words, you remove duplicates one at a time. Don't kill the messager; learn from the message. If there is one area within the data quality space that Microsoft needs to improve on, it's this one.
That said, let's move on. So you've detected duplicates and now you want to eliminate them from your data.
If you remember from last month's post, read up here if you don't, a duplicate detection job returns potential duplicates and allows you to browse each one along with it's potential match. Consult the screenshot below for a view of what that looks like.
[caption id="attachment_306" align="aligncenter" width="500" caption="Duplicate Detection Job results"]
In the lower pane of the screenshot above there is a toolbar option (3rd from the left) is a icon to merge the two highlighted records. This is where the consolidation effort begins.
One of the best features of the merge functionality is that it has the flexibility to build a composite, or best of available information, master record. Briefly, the master record is the record which is retained as the active record. It also allows the end user to select one record over another. An example of these features is outlined in the screenshot below. First let's look at the option where each element of the master record are selected.
[caption id="attachment_312" align="alignleft" width="500" caption="An example of an all inclusive master record selection"]
Here's a look at the composite option:
[caption id="attachment_315" align="aligncenter" width="500" caption="An example of a composite master record selection "]
Notice in the all inclusive example the entire left hand column is highlighted in blue, whereas in the composite option only those elements selected via the radio button are highlighted in blue. This is a visual indication of what data elements will be retained in the master record.
This is one of my favorite pieces of functionality with regard to the merge option. Often end users vary in the data they provide and it is always better for an organization to retain as much information about their customers as is possible.
I specifically chose the composite screenshot presented because it illustrates one of those important aspects in customer data quality. Noticed that the element selected from the right hand side was a middle initial. This data element is invaluable when performing data matching and having that element can make an important distinction between two different customers later on down the road.
Once you've defined what your master record looks like, either via the all inclusive or composite method, it is time to commit that selection. The screenshot below illustrates how this is performed.
[caption id="attachment_316" align="aligncenter" width="500" caption="How to commit your master record selection"]
An important option in the commit process is one enabled through the checkbox provided visible in the screenshot above. Not every field in a record is in all cases exposed via the merge utility. As a result the option made available through the checkbox allows you to make sure that, as the label indicates, select every field with data from the chosen master record even if there is a different value in the other record. Simply put, it is an overwrite function that retains all the data from the selected master record beyond what is visible in the merge screen.
Once you've reviewed and are confident in your selection, you simply click on the OK button. Provided there are no commit locks on the record, which indicates that another user has one of the two records open and is actively working on it, you will receive the following dialog box confirming your consolidation success.
[caption id="attachment_317" align="aligncenter" width="441" caption="Duplicate elimination success!"]
It is critical to note that the subordinate, or non-master record, is NOT deleted from the system. It is simply deactivated. This is to say that a flag (statecode) is changed to inactivate. One important note about the statecode field is that unlike conventional notation a value of '1' is not active in Microsoft Dyanmics CRM. Instead Microsoft chose the value '0' as active and '1' as inactive. Consequently all non-master records in CRM have a statecode value of '1'. This little fact can save hours of data analysis and perserve the samity of your DBAs, so it is worth noting.
I hope this information was beneficial to you Microsoft Dyanmics CRM users and administrators. As usual I welcome all comments, questions, and suggestions. So please feel free to comment on this post and I'll try and replay in a timely manner.
Tuesday, September 1, 2009
August Edition of IAIDQ Festival del IDQ Bloggers
I am glad to be hosting the Festival del IDQ Bloggers this month! I've tried to capture the core of each message, but each of these is worth a deeper look. Don't forget to follow the submission links and get all the details!
This month's first submission comes to use from Daragh O Brien. Daragh poses an interesting question when he asks, Is information quality management a recession proof profession?
One clear take away from this post is that those who express the value proposition of an information quality initiative are more likely to be regarded as valuable and necessary. In that way those who participate in these initiatives can be thought of a recession proof.
About Daragh: Obsessive blogger, information quality consultant and Director of IAIDQ with over 12 years experience at the sharp end of Information Quality. Taoiseach (CEO) of Castlebridge Associates, a specialist Information Quality consulting business based in Ireland.
Continuing on a theme our second post comes to us from Dylan Jones and shows us How To Deliver A Compelling Data Quality Business Case.
Dylan recommends several excellent ways to build and deliver the value proposition such as:
- use time/date stamps to show information quality as a long term problem and not a short lived glitch
- use cause and affect analysis to link data quality issues with business process gaps
- review the annual report to gain insight into corporate strategies that can benefit from information quality services
- don't use PowerPoint
- keep it simple to get the point across
About Dylan: Dylan is the founder and editor of the Data Quality Pro which is dedicated to "helping data quality professionals take their career or business to the next level."
Switching gears a little, Jim Harris reminds us to keep our theories in check until after we've taken the time to really listen to what's being said in his post Hailing Frequencies Open.
In this post Jim points out the difference between waiting to talk and what is called empathetic listening where we are actually listening with the intent to really try to understand the other person's frame of reference.
About Jim: Jim Harris is an independent consultant, speaker, writer and blogger with over 15 years of professional services and application development experience in data quality. Obsessive-Compulsive Data Quality is an independent blog offering a vendor-neutral perspective on data quality.
While we are on the topic of communication Steve Sarsfield recommends some data governance dialog for CEO's in his post, 9 Questions CEOs Should Ask About Data Governance.
In this post Steve points out that the executive team is responsible to lower risk and gain control through their influence on data governance. The following are a few of the question Steve suggestions:
- Do we have a data management strategy?
- Are we in compliance with all laws regarding our governance of data?
- Do you have the access to data you need?
About Steve: Steve Sarsfield is a data quality evangelist and author of the book the Data Governance Imperative.
Finally in Sweden meets United States a post from Henrik Liliendahl Sørensen on data matching and how character sets, address formats and naming conventions are just a few of the complexities when the data originates in a different language. Maybe, as he suggests, centralized reference data is a step in the right direction to solving some of these issues?
About Henrik: Henrik is a Data Quality and Master Data Management professional also doing Data Architecture. You can check out what is on his mind at Liliendahl on Data Quality.
I hope you enjoyed this month's edition of the blog carnival! I want to thank all those who've submitted postings and encourage those who have not done so yet to participate in this opportunity.
Tuesday, August 4, 2009
The DQ Two Step!
Heck! You can even grab your partner if the mood strikes you?
[caption id="attachment_203" align="alignleft" width="85" caption="Feels good!"]
I've recently cleansed some customer data for a great bunch of folks that I love calling my client.
We had a consolidation ratio of 1:6 or around 17%. Which equated to roughly a million duplicates removed. That's a lot savings on postage stamps for the marketing department so they were psyched! We validated over 90% of the addresses and built reports to identify those that did not meet the requirements for a valid address. Not too shabby if I don't say so myself!
Now that we've deployed the data into User Acceptance (UAT) I find myself in a familiar place; the business logic.
[caption id="attachment_202" align="alignleft" width="100" caption="You missed something"]
You can spend all the time you need, or even care to, on rules for consolidation but it usually is not until the data hits the screen that the ramifications are easily understood by the average business user.
Case in point, I recently received an email from a stakeholder asking me to look over some data with him. I was curious what I'd find when I reached his office as I analyzed this data and the processing code more than a few times by now. On my walk over I went through many possible scenarios in my head.
Was it something I missed? Surely not. I've performed several test runs in order to validate the business logic. With my curiosity peaked I rounded the corner and into his office I went.
[caption id="attachment_204" align="alignleft" width="150" caption=""Good point!""]
After a little chit-chat, like I said I love this client, we got down to business. He proceeded to type a few parameters in the search utility and I waited with anticipation.
However after a second, maybe less, my anticipation was replaced with relief and more than a little disbelief. I'd been over this a time or two which is why I was in a state of shock. Not to mention my client was not someone who needed "Data for Dummies".
With identities masked to protect the innocent, below is a sample of the records he was concerned about and wanted me to see.
[caption id="attachment_208" align="aligncenter" width="468" caption="Duplicate?"]
[caption id="attachment_211" align="aligncenter" width="468" caption="They've been through this before!"]
What data quality is (and what it is not)
Like the radar system pictured above, data quality is a sentinel; a detection system put in place to warn of threats to valuable assets. ...
-
Recently I had coffee with Dr. John Talburt of the University of Arkansas at Little Rock's Information Quality program . During the conv...
-
[caption id="" align="alignleft" width="240"] Image by Marius B via Flickr[/caption] Is Big Data better Data ...