Showing posts with label master data management. Show all posts
Showing posts with label master data management. Show all posts

Tuesday, October 16, 2012

Brief Thoughts on Data Virtualization and Master Data Management


It was nice to see, in a recent book I have been reading, some recognition of the usefulness of master data management (MDM), and how the functions included in data virtualization solutions give users a flexibility in architecture design that’s worth its weight in gold (IBM and other vendors, take note). What I have perhaps not sufficiently appreciated in the past is its usefulness in speeding MDM implementation, as duly noted by the book.

I think that this is because I have assumed that users will inevitably reinvent the wheel and replicate the functions of data virtualization in order to afford users a spectrum of choices between putting all master data in a central database and only there, and leaving the existing master data right where it is.  It now appears that they have been slow to do so.  And that, in turn, means that the “cache” of a data virtualization solution can act as that master-data central repository while preserving the farflung local data that compose it. Or, the data virtualization server can provide discovery of the components of a customer master-data record, give a sandbox to define a master data record, alert to new data types that will need to change the master record, and enforce consistency – all key functions of such a flexible MDM solution.

But the major value-add is the speedup of implementing the MDM solution in the first place, by speeding definition of master data, writing code on top of it for application interfaces, and allowing rapid but safe testing and upgrade. As the book says, abstraction gives these benefits.

Therefore, it continues to be worth it for both existing and new MDM implementations to seriously consider data virtualization.

Wednesday, October 10, 2012

Data Virtualization: Users Grok What I Said 10 Years Ago


Ten years ago I put out the first EII (now data virtualization) report.  In it I said:
  • ·         The value of DV is both in its being a database veneer across disparate databases, and in its discovery and storage of enterprise-wide global metadata
  • ·         DV can be used for querying and updates
  • ·         DV is of potential value to users and developers and administrators. Users see a far wider array of data, and new data sources are added more quickly/semi-automatically. Developers have “one database API to choke.” Administrators can use it to manage multiple databases/data stores at a time, for cost savings
  • ·         DV can be great for mergers, metadata standardization, and as a complement to existing enterprise information architectures (obviously, including data warehousing)
  • ·         DV is a “Swiss army knife” that can be used in any database-related IT project
  • ·         DV is strategic, with effects not just on IT costs but also on corporate flexibility and speed to implement new products (I didn’t have the concept of business agility then)
  • ·         DV can give you better access to information outside the business.
  •  DV can serve as the “glue” of an enterprise information architecture.
I’m at DV Day 2012 (Composite Software) in NYC. Today, for the first time, I have heard not just vendors talking about implementing these things, but users actually doing them – global metadata, user self-service access to the full range of corporate data, development using SQL access, use for updating in operational databases, use for administration of at least the metadata management of multiple databases as well as archiving administrative tasks, use in mergers, use in “data standardization”, use as an equal partner with data warehousing, use in just about every database-related IT project, selling by as strategic to the business with citations of impacts on the bottom line through strategic products as well as “business agility”, use to access extra-enterprise Big Data across multiple clouds, and use as the framework of an enterprise information architecture.

I just wanted to say, on a personal note, that this is what makes being an analyst worthwhile.  Not that I was right 10 years ago. That, for once, by the efforts of everyone pulling together to make the point and elaborate the details, we have managed to get the world to implement a wonderful idea. I firmly believed that, at least in some small way, the world would be better off if we implemented DV. Now, it’s clear that that’s true, and that DV is unstoppable.

And so, I just wanted to take a moment to celebrate.

Sunday, September 23, 2012

Is Your Organization Suffering From Data Warehouse Disease?


I have a feeling that a fair amount of readers – especially vendors and IT BI types – are going to be upset by what I have to say in this post.  However, viewing some of the material that has passed across my desk recently, I really think it’s time to raise the question of whether too much organizational power given to data warehouse folks is beginning to cause some significant under-performance in meeting today’s key organizational information management needs.

The immediate occasion for these reflections is that I am partway through a book on a related subject that goes into some detail on data warehousing’s view of the world:  how BI should be handled, what the organizational information architecture should be, and how we got this way.  This book will remain nameless, because in many ways it’s an excellent primer.  However, over the last 22-31 years (depending on whether you count my software development days), I have had a cross-organization, cross-vendor view of the same area, and I have to say that the book redefines history and the purposes of various things in the ideal information architecture in major ways.

Usually, I find that going over history just wastes time in a blog post – but here, it helps to see how data warehouse concepts of common information management terms make them reinterpret the purposes of the underlying products, making the information architecture – and the whole information handling process – potentially (and, probably, actually) less effective in the medium and long term. So let’s combine history and exposition of my assertion.

A Data Warehousing View of the World

In brief, the book’s view of the information architecture seems to be as follows: Data of all types comes in to production systems, which immediately pass it on to the data warehouse for cleansing and aggregation. Behind the data warehouse is an optional operational data store for key data, and things like master data management operate in parallel with the data warehouse to provide a global view of multiple local ways to store customer data. On top of the data warehouse are key Business Intelligence applications, which include both repetitive, scheduled reporting and analytics.

Now, this view of the world seems reasonable if you were born yesterday, or if you’ve spent the last fifteen years entirely in data warehousing.  However, there are, in my view, some major problems with it.
In the first place, afaik, only in data warehousing are the databases at the initial entry point referred to as “production systems”.  For twenty years, I have been calling them “operational databases”. In fact, they were business-critical before data warehousing existed, and so were the apps on top of them – like ERP. 

Why does this matter? Because it allows data warehouse folks to shift the “operational data store” behind the data warehouse.  The operational data store is a later concept, and one that I (among others, I assume) wrote papers proposing around 2004 and 2005. The idea is that the data warehouse is simply too slow to react immediately to key operational data – but that operational data is scattered across multiple operational data stores, and so an “operational data store” makes sure that a subset of operational data for quick decision-making is either put in a central point for quick analysis in parallel with its arrival, or monitored by a central “virtual database.” Putting the operational data store behind the data warehouse defeats its entire purpose.

Likewise, the master data management system. I wrote papers on this in assessing IBM’s version of the concept in 2006 and 2007. Again, the notion was of combining operational data coming in to operational databases – in this case, by enforcing a common format that allowed cross-organization and cross-country leveraging of operational data by ERP and customer intelligence apps. By redefining the master data management as existing within the data warehouse or at the same remove from operational databases, data warehouse folks ensure that master data management moves no faster than the data warehouse.

And finally, there is the idea that (implicitly) analytics is entirely contained in BI, and hence is entirely dependent on the data warehouse. On the contrary, an increasing amount of analytics goes on outside of BI.  For example, analytics is part of products that analyze computer infrastructure semi-automatically to optimize performance or detect upcoming problems. Or, it is used to analyze key computer-supported business processes.  This is “intelligence” in the sense of “military intelligence” – proactively going out and finding out what’s going on – but it is not “business intelligence” in the sense of finding out what’s going on inside and outside the business on the basis of data that is handed to you, and that your reporting tools are too slow or shallow to tell you. In other words, these applications of analytics are entirely outside of a reactive data warehouse.

Why It Matters

There are two places that over-emphasis on data warehousing can impede organizational BI and other information management effectiveness:  the information architecture, and the organization’s “agility” in responding to new kinds of information from outside. As I’ve suggested in the previous section, a data warehousing view of the information architecture shifts operations that involve lots of “updates” and data just arrived from outside to the data warehouse or behind it.  That means going through the data-warehouse cleansing and aggregation process and arriving in a centralized location that is handling queries from all over the organization and is optimized for adding new data not “on the fly” but in delayed bursts. There is simply no way that is going to be as timely as performing tasks on the data as it arrives in the operational systems.

Just as troubling, the entire emphasis of the organization is now more reactive and focused farther away from the organization’s “antennae” to the outside environment. The IT organization appears to be focused on responding to new demands from business for timelier data, not actively seeking the latest new information and merging it back into existing systems. The IT organization appears to emphasize cleaning up the data and merging it and only then analyzing it at an internal “choke point”, rather than handling the information faster where it arrives. 

If you think these concerns are theoretical, think about the case of social-media Big Data. Yes, Oracle as a major vendor is emphasizing inhaling huge amounts of this data from multiple clouds into the data warehouse and then analyzing it – when the whole purpose of the NoSQL movement is to allow rapid in-cloud analysis of inconsistent, uncleansed data – but it would not do so unless there was some organizational push to avoid analytics outside the data warehouse.  I conclude that there is some strong evidence that a data warehousing focus is impeding organizational ability to process and feed to business decision makers key information in as timely a fashion as possible.

Moreover, there is some sense that this is not an organizational quirk but a tendency so embedded in the IT organization that this impediment is a symptom not of a temporary problem that is easy to fix, but rather of an organizational “disease.” In other words, simply directing the organization to pay more attention to doing social-media processing in the cloud will probably not work.

Action Strategies and Conclusion

First (although I think there is little danger of this) I must caution against throwing the baby out with the bathwater.  There are very good reasons to have a data warehouse performing the core functions of querying for BI. I have, in the past, conjectured that if I were to design a new information architecture today, I might not create a data warehouse or data mart at all – instead, I might impose “data virtualization” and master data management tools over existing operational databases. However, practically speaking, in most if not all cases, the sheer experience behind today’s data warehousing products makes them far more preferable for core functions.

Rather, I would suggest that data warehousing be placed under, and be responsive to rather than dominant over, an information architecture and information strategy function aimed more at the edge of the organization than its central data center. This is not a matter of making the organization more responsive to the business; it is a matter of making the IT organization more agile (by my definition, which stresses the utility of proactive and outside-the-organization-directed agility).

Until I saw this book, which suggested that data warehouse folks had gone too far in asserting “IT information handling is all about the data warehouse”, I was not too concerned about data warehousing folks; I would get into annoying arguments with folks who thought I just didn’t “get” data warehousing, but it seemed to me that the benefits of a powerful database-related IT function outweighed the negatives of data warehouse folks’ “not invented here” blind spots.  Now, I am rethinking my position.  If the result of this type of rewriting of history is an increasingly sub-optimal information architecture, then such a “disease” is not so harmless after all.

Does your organization suffer from data warehouse disease?  If so, what do you think should be done about it?

Friday, March 16, 2012

Information Agility and Data Virtualization: A Thought Experiment

In an agile world, what is the value of data virtualization?

For a decade, I have been touting the real-world value-add of what is now called data virtualization – the ability to see disparate data stores as one, and to perform real-time data processing, reporting, and analytics on the “global” data store – in the traditional world of IT and un-agile enterprises. However, in recent studies of agile methodologies within software development and in other areas of the enterprise, I have found that technologies and solutions well-adapted to deliver risk management and new product development in the traditional short-run-cost/revenue enterprises of the past can actually inhibit the agility and therefore the performance of an enterprise seeking the added advantage that true agility brings. In other words: your six-sigma quality-above-all approach to new product development may seem as if it’s doing good things to margins and the bottom line, but it’s preventing you from implementing an agile process that delivers far more.

In a fully agile firm – and most organizations today are far from that – just about everything seems mostly the same in name and function, but the CEO or CIO is seeing it through different eyes. The relative values of people, processes, and functions change, depending no longer on their value in delivering added profitability (that’s just the result), but rather on the speed and effectiveness with which they evolve (“change”) in a delicate dance with the evolution of customer needs.

What follows is a conceptualization of how one example of such a technology – data virtualization – might play a role in the agile firm’s IT. It is not that theoretical – it is based, among other things, on findings (that I have been allowed by my employer at the time, Aberdeen Group, to cite) about the agility of the information-handling of the typical large-scale organization. I find that, indeed, if anything, data virtualization – and here I must emphasize, if it is properly employed – plays a more vital role in the agile organization.

Information Agility Matters

One of the subtle ways that the ascendance of software in the typical organization has played out is that, wittingly or unwittingly, the typical firm is selling not just applications (solutions) but also information. By this, I don’t just mean selling data about customers, the revenues from which the Facebooks of the world fatten themselves on. I also mean using information to attract, sell, and follow-on sell customers, as well as selling information itself (product comparisons) to customers.

My usual example of this is Amazon’s ability to note customer preferences for certain books/music/authors/performers, and to use that information to pre-inform customers of upcoming releases, adding sales and increasing customer loyalty. The key characteristic of this type of information-based sale is knowledge of the customer that no one else can have, because it involves their sales interactions with you. And that, in turn, means that the information-led sell is on the increase, because the shelf life of application comparative advantage is getting shorter and shorter – but comparative advantage via proprietary information, properly nurtured, is practically unassailable. In the long run, applications keep you level with your competitors; information lets you carve out niches in which you are permanently ahead.

This, however, is the traditional world. What about the agile world? Well, to start with, the studies I cited earlier show that in almost any organization, information handling is not only almost comically un-agile but also almost comically ineffective. Think of information handling as an assembly line (yes, I know, but here it fits). More or less, the steps are:

1. Inhale the data into the organization (input)

2. Relate it to other data, including historical data, so the information in it is better understood (contextualize)

3. Send it into data stores so it is potentially available across the organization if needed (globalize)

4. Identify the appropriate people to whom this pre-contextualized information is potentially valuable, now and in the future, and send it as appropriate (target)

5. Display the information to the user in the proper context, including both corporate imperatives and his/her particular needs (customize)

6. Support ad-hoc additional end-user analysis, including non-information-system considerations (plan)

The agile information-handling process, at the least, needs to add one more task:

7. Constantly seek out new types of data outside the organization and use those new data types to drive changes in the information-handling process – as well, of course, as in the information products and the agile new-information-product-development processes (evolve).

The understandable but sad fact is that in traditional IT and business information-handling, information “leaks” and is lost at every stage of the process – not seriously if we only consider any one of the stages, but (by self-reports that seem reasonable) yielding loss of more than 2/3 of actionable information by the end of steps one through six. And “fixing” any one stage completely will only improve matters by about 11% -- after which, as new types of data arrive, that stage typically begins leaking again.

Now add agile task 7, and the situation is far worse. By self-reports three years ago (my sense is that things have only marginally improved), users inside an organization typically see important new types of data on average ½ year or more after they surface on the Web. Considering the demonstrated importance of things like social-media data to the average organization, as far as I’m concerned, that’s about the same as saying that one-half the information remaining after steps one to six lacks the up-to-date context to be truly useful. And so, the information leading to effective action that is supplied to the user is now about 17% of the information that should have been given to that user.

In other words, task 7 is the most important task of all. Stay agilely on top of the Web information about your customers or key changes in your regulatory, economic, and market environment, and it will be almost inevitable that you will find ways to improve your ability to get that information to the right users in the right way – as today’s agile ad-hoc analytics users (so-called business analysts) are suggesting is at least possible. Fail to deliver on task 7, and the “information gap” between you and your customers will remain, while your information-handling process falls further and further out of sync with the new mix of data, only periodically catching up again, and meanwhile doubling down on the wrong infrastructure. Ever wonder why it’s so hard to move your data to a public cloud? It’s not just the security concerns; it’s all that sunk cost in the corporate data warehouse.

In other words, in information agility, the truly agile get richer, and the not-quite-agile get poorer. Yes, information agility matters.

And So, Data Virtualization Matters

It is one of the oddities of the agile world that “evolution tools” such as more agile software development tools are not less valuable in a methodology that stresses “people over processes”, but more valuable. The reason why is encapsulated in an old saying: “lead, follow, or get out of the way.” In a world where tools should never lead, and tools that get out of the way are useless, evolution tools that follow the lead of the person doing the agile developing are priceless. That is, short of waving a magic wand, they are the fastest, easiest way for the developer to get from point A to defined point B to unexpected point C to newly discovered point D, etc. So one should never regard a tool such as Data Virtualization as “agile”, strictly speaking, but rather as the best way to support an agile process. Of course, if most of the time a tool is used that way, I’m fine with calling it “agile.”

Now let’s consider data virtualization as part of an agile “new information product development” process. In this process, the corporation’s traditional six-step information-handling process is the substructure (at least for now) of a business process aimed at fostering not just better reaction to information but also better creation and fine-tuning of combined app/information “products.”

Data virtualization has three key time-tested capabilities that can help in this process. First is auto-discovery. From the beginning, one of the nice side benefits of data virtualization has been that it will crawl your existing data stores – or, if instructed, the Web – and find new data sources and data types, fit them into an overall spectrum of data types (from unstructured to structured), and represent them abstractly in a metadata repository. In other words, data virtualization is a good place to start proactively searching out key new information as it sprouts outside “organizational boundaries”, because they have the longest experience and the best “best practices” at doing just that.

Data virtualization’s second key agile-tool capability is global contextualization. That is, data virtualization is not a cure-all for understanding all the relationships between an arriving piece of data and the data already in those data stores; but it does provide the most feasible way of pre-contextualizing today’s data. A metadata repository has been proven time and again to be the best real-world compromise between over-abstraction and zero context. A global metadata repository can handle any data type thrown at it. For global contextualization, a global metadata repository that has in it an abstraction of all your present-day key information is your most agile tool. It does have some remaining problems with evolving existing data types; but that’s a discussion for another day, not important until our information agility reaches a certain stage that is still far ahead.

Data virtualization’s third key agile-tool capability is the “database veneer.” This means that it allows end users, applications, and (although this part has not been used much) administrators to act as if all enterprise data stores were one gigantic data store, with near-real-time data access to any piece of data. It is a truism in agile development that the more high-level the covers-all-cases agile software on which you build, the more rapidly and effectively you can deliver added value. The database veneer that covers all data types, including ones that are added over time, means more agile development of both information and application products on top of the veneer. Again, as with the other two characteristics, data virtualization is not the only tool out there to do this; it’s just the one with a lot of experience and best practices to start with.

And, for some readers, that will bring up the subject of master data management, or MDM. In some agile areas, MDM is the logical extension of data virtualization, because it begins to tackle the problems of evolving old data types and integrating new ones in a particular use case (typically, customer data) – which is why it offers effectively identical global contextualization and database-veneer capabilities. However, because it does not typically have auto-discovery and is usually focused on a particular use case, it isn’t as good an agile-tool starting point as data virtualization. In agile information-product evolution, the less agile tool should follow the lead of the more agile one – or, to put it another way, a simple agile tool that covers all cases trumps an equivalently simple agile tool that only covers some.

So, let’s wrap it up: You can use data virtualization, today, in a more agile fashion than other available tools, as the basis of handling task 7, i.e., use it to auto-discover new data types out on the Web and integrate them continuously with the traditional information-handling process via global contextualization and the database veneer. This, in turn, will create pressure on the traditional process to become more agile, as IT tries to figure out how to “get ahead of the game” rather than drowning in a continuously-piling-up backlog of new-data-type requests. Of course, this approach doesn’t handle steps 4, 5, and 6 (target, customize, and plan); for those, you need to look to other types of tools – such as (if it ever grows up!) so-called agile BI. But data virtualization and its successors operate directly on task 7, directly following the lead of an agile information-product-evolution process. And so, in the long run, since information agility matters a lot, so does data virtualization.

Some IT-Buyer Considerations About Data Virtualization Vendors

To bring this discussion briefly down to a more practical level, if firms are really serious about business agility, then they should consider the present merits of various data virtualization vendors. I won’t favor one against another, except to say that I think that today’s smaller vendors – Composite Software, Denodo, and possibly Informatica – should be considered first, for one fundamental reason: they have survived over the last nine years in the face of competition from IBM among others.

The only way they could have done that, as far as I can see, is to constantly stay one step ahead in the kinds of data types that customers wanted them to support effectively. In other words, they were more agile – not necessarily very agile, but more agile than an IBM or an Oracle or a Sybase in this particular area. Maybe large vendors’ hype about improved agility will prove to be more than hype; but, as far as I’m concerned, they have to prove it to IT first, and not just in the features of their products, but also in the agility of their data-virtualization product development.

The Agile Buyer’s Not So Bottom Line

Sorry for the confusing heading, but it’s another opportunity to remind readers that business agility applies to everything, not just software development. Wherever agility goes, it is likely that speed and effectiveness of customer-coordinated change are the goals, and improving top and bottom lines the happy side-effect. And so with agile buyers of agile data-virtualization tools for information agility – look at the tool’s agility first.

To recap: the average organization badly needs information agility. Information agility, as in other areas, needs tools that “follow” the lead of the person(s) driving change. Data virtualization tools are best suited for that purpose as of now, both because they are best able to form the core of a lead-following information-product-evolution toolset, and because they integrate effectively with traditional information-handling and therefore speed its transformation into part of the agile process. Today’s smaller firms, such as Composite Software and Denodo, appear to be farther along the path to lead-following than other alternatives.

Above all: just because we have been focused on software development and related organizational functions, that doesn’t mean that information agility isn’t equally important. Isn’t it about time that you did some thought experiments of your own about information agility? And about data virtualization’s role in information agility?