Showing posts with label BI. Show all posts
Showing posts with label BI. Show all posts

Tuesday, February 12, 2013

GoodData and Grey's Anatomy: It's Not Always the Vendor's Fault


Recently I received information about an interesting new vendor called GoodData – it appears that they are specializing in BI-type user interfaces for function-specific dashboards (of course, in the cloud and involving cross-database information display).  One of these was GoodSupport Bash, a “bash-up” of customer-support metrics and KPIs (Key Performance Indicators).  And that caused me to reflect on the idea of applying these types of “comprehensive” metrics to customer support.

Customer-support technology has a relatively long software-vendor history:  I remember first seeing the idea in action back in 1982, when Computer Corp. of America’s email system was used to track customer inquiries. Our forms capability allowed us to enter the call info in an email, send it for resolution to the appropriate techie, and then see the “email trail” as the problem was resolved.  Like many small companies, we had a reputation for excellent customer support because the techies were developers who had spare time to actually solve problems and the clout to change the code base to do so.  In fact, I was once told that one “guru” would go to install at new sales and when the customer asked for changes, he would make them without documentation or changing the main code base, so that the next version overrode the customization – it made the customer very happy initially, and very annoyed a year later.

However, in the early 2000s, when offshoring came in, major-vendor customer support was one of the first to move, and the semi-mature customer-support software of the day now had to define support personnel tasks very carefully, as the new Indian and Filipino respondents had some difficulties with American English and more difficulties with being up to date with software technology.  The resultant upheaval in customer support has never really settled down again, imho.  Companies like Microsoft would really like to move away from live-person customer support entirely, but the complexity of the technology means that users keep having to ask questions of live representatives who can be forced to pay attention.  For example, it was not until I got to a live representative that I realized that the Windows 8 version of Windows Media Player does not support DVD playing out of the box (according to PC World, you need the $10 Media Pro or some such).

Well, this in turn meant that phone-based customer support had to be “optimized” in some way, else waits of 2 hours and unresponsive representatives would cause a significant black eye for the company providing the support for its products.  And so, solutions like GoodSupport Bash allow companies to “take the temperature” of their customer support according to all sorts of criteria, like speed to reach a representative, time taken by the representative, and whether the representative is taking the opportunity for up-sell or cross-sell.

Except that I believe that the results of this kind of fine-tuning according to various management theories can sometimes actually be counter-productive – and the fault lies not with the customer-support software vendor.

Grey’s Anatomy and Ice Removal

I must confess that Grey’s Anatomy is one of my guilty pleasures, and I often find it difficult to watch without rotfl.  To me, GA is the ultimate in actor torture.  In just about every episode, at least one and often several actors and actresses must somehow make credible a complete 360-degree turn in their characters, with the words they are given to do this stretching the limits of belief that anyone could talk in this way.  As character after character changes partners or sexual orientation on the verge of a marriage and in the middle of a surgery, I can almost imagine the reaction as that actor or actress sees the week’s script – bracing themselves, and yet unable to anticipate the next contortion.  Although there are limits:  no one has taken a sudden interest in animal husbandry – yet.

Anyway, in this episode, doctors facing a hostile takeover of the hospital went to find out what this new highly-efficient corporation was doing, and posed as patients.  The doctor started reading from a scripted set of questions, and when the “patients” attempted to derail them, started the script over again until they could get it done.  This was done, we are assured, in the name of standardized efficiency.  You will recognize the similarity to some customer support experiences.

However unlikely the situation, the GA actors and actresses, trained to make the unlikely plausible, made it very clear how customer support driven by metrics appears to the customer.  The caller reporting something that often is unusual must be fit into a series of questions that attempts to cover only the usual.  Callers who do not fit are effectively ignored – you could share the frustration as the “patient” attempting to vaguely describe symptoms is asked to spend much time talking about things that are not problems, and the feeling that no one is listening.  At one point, the “patient” confesses that he is a doctor, pretends he’s already a new part of the company, and asks why these procedures; the “representative” assures him in a scripted manner that the procedures are more effective – except it’s clear that he must answer that way, no matter what he thinks, or lose his job.

Sound implausible?  Well, I actually went through a similar experience the other day, when I tried to see if a national gutter specialist could provide snow and ice removal on an emergency basis.  Sure enough, I wound up with a representative who took some time to realize that I was not calling about gutter cleaning, finally said that the company did not provide the services, and then spent three minutes attempting to upsell me on a bigger gutter cleaning contract, despite the fact that (a) I said twice that I was not going to do so, and (b) I had just been told that they didn’t do something else I needed.  It was clear that this was one of the metrics by which he was judged:  did you run through all the questions?  Did you upsell? You can imagine what I felt about the company afterwards – and it wasn’t the representative’s fault.

Sloan Management Review, iirc, had some interesting data on the effect of this type of thing on customer loyalty.  Often, the most loyal or biggest-spending before the disappointment were the quickest to jump ship – they felt “entitled”, as it were.  But even the truly loyal needed customization that truly paid attention to their needs, else they too found it difficult to stay.  Branding only goes so far.

My point here, however, is that none of this was the customer-support software vendor’s fault.  On the contrary, whether the company’s support strategy is good or bad, the typical indication from users is that something like GoodData Support Bash will make it more efficient.  If the strategy is bad, however, it may also make it less effective, or more harmful.

The Bottom Line:  Metrics Is As Metrics Does

It is now a truism of management theory that metrics create behavior – employees game everything to their own advantage, or to minimize their own disadvantage.  And so, users of customer-support software like GoodData Support Bash badly need to remember that the most effective use of any customer-facing tool is not to create representatives who will deliver the company’s idea of customer happiness via efficiency, but to find analytics to understand the customer better.  Only after that knowledge is gained and the company has used that knowledge to improve customer support responsiveness to real customer needs should efficiency metrics be considered.

To put it more concretely:  among the metrics in your dashboard, is at least one telling you honestly the degree of customer satisfaction with the interaction?  And are you mining data from that metric telling you what’s good and bad about your present solutions?  If not, are you really making things better with your improvements in customer-support efficiency?

Or are you making them worse?

Monday, October 15, 2012

Does Data Virtualization Cause a Performance Hit? Maybe The Opposite

In the interest of “truthiness”, consultants at Composite Software’s Data Virtualization Day last Wednesday said that the one likely tradeoff for all the virtues of a data virtualization (DV) server was a decrease in performance.  It was inevitable, they said, because DV inserts an extra layer of software between a SQL query and the engine performing that query, and that extra layer’s tasks necessarily increases response time. And, after all, these are consultants who have seen the effects of data-virtualization implementation in the real world. Moreover, it has always been one of the tricks of my analyst trade to realize that emulation (in many ways, a software technique similar to DV) must inevitably involve a performance hit – due to the additional layer of software.

And yet, I would assert that three factors difficult to discern in the heat of implementation may make data virtualization actually better-performing than the system as it was pre-implementation.  These are:
·         Querying optimization built into the data virtualization server

·         The increasingly prevalent option of cross-data-source queries

·         Data virtualization’s ability to coordinate across multiple instances of the same database.
Let’s take these one at a time.

DV Per-Database Optimization


I remember that shortly after IBM released their DV product in around 2005, they did a study in which they asked a bunch of their expert programmers to write a set of queries against an instance of IBM DB2, iirc, and then compared their product’s performance against these programmers’.  Astonishingly, their product won – and yet, seemingly, the deck was entirely stacked against it.  This was new programmer code optimized for the latest release of DB2, based on in-depth experience, and the DV product had that extra layer of software. What happened?
According to IBM, it was simply the fact that no single programmer could put together the kind of sophisticated optimization that was in the DV product. Among other things, this DV optimization considered not only the needs of the individual querying program, but also its context as part of a set of programs accessing the same database.  Now consider that in the typical implementation, the deck is not as stacked against DV: the programs being superseded may have been optimized for a previous release and never adequately upgraded, or the programmers who wrote them or kept them current with the latest release may have been inexperienced.  All in all, there is a significant chance (I wouldn't be surprised if it was better than 50%) that DV will perform better than the status quo for existing single-database-using apps “out of the box.”
Moreover, that chance increases steadily over time – and so an apparent performance hit on initial DV implementation will inevitably turn into a performance advantage 1-2 years down the line. Not only does the percentage of “older” SQL-involving code increase over time, but the need for database upgrade (as, for example, upgrading DB2 every 2 years should pay off in spades, according to my recent analyses), means that these DV performance advantages widen – and, if emulation is any guide, the performance cost from an extra layer never gets worse than 10-20%.

The Cross-Data-Source Querying Option


Suppose you had to merge two data warehouses or customer-facing apps as part of a takeover or merger.  If you used DV to do so, you might see (as in the previous section) the initial queries to either app or data warehouse be slower.  However, it seems to me that’s not the appropriate comparison.  You have to merge the two somehow.  The alternative is to physically merge the data stores and maybe the databases accessing those data stores.  If so, the comparison is with a merged data store for which neither set of querying code is optimized, and a database for which one of the two sets of querying code, at the least, is not optimized. In that case, DV should have an actual performance advantage, since it provides an ability to tap into the optimizations of both databases instead of sub-optimizing one or both.
And we haven’t even considered the physical effort and time of merging two data stores and two databases (very possibly, including the operational databases, many more than that).  DV has always sold itself on its major advantages in rapid implementation of merging – and has constantly proved its case.  It is no exaggeration to say that a year saved in merger time is a year of database performance improvement gained.
Again, as noted above, this is not obvious in the first DV implementation.  However, for those who care to look, it is definitely a real performance advantage a year down the line.
But the key point about this performance advantage of DV solutions is that this type of coordination of multiple databases/data stores instead of combining them into one or even feeding copies into one central data warehouse is becoming a major use case and strategic direction in large-enterprise shops. It was clear from DV Day that major IT shops have finally accepted that not all data can be funneled into a data warehouse, and that the trend is indeed in the opposite direction.  Thus, an increasing proportion (I would venture to say, in many cases approaching 50%) of corporate in-house data is going to involve cross-data-source querying, as in the merger case. And there, as we have seen, the performance advantages are probably on the DV side, compared to physical merging.

DV Multiple-Instance Optimization


This is perhaps a consideration more suited to abroad and to medium-sized businesses, where per-state or regional databases and/or data marts must be coordinated. However, it may well be a future direction for data warehouse and app performance optimization – see my thoughts on the Olympic database in a previous post. The idea is that these databases have multiple distributed copies of data. These copies have “grown like Topsy”, on an ad-hoc, as needed basis.  There is no overall mechanism for deciding how many copies to create in which instances, and how to load balance across copies.
That’s what a data virtualization server can provide.  It automagically decides how to optimize given today’s incidence of copies, and ensures in a distributed environment that the processing is “pushed down” to the right database instance. In other words, it is very likely that data virtualization provides a central processing software layer rather than local ones – so no performance hit in most cases – plus load balancing and visibility into the distribution of copies, which allows database administrators to achieve further optimization by changing that copy distribution. And this means that DV should, effectively implemented, deliver better performance than existing solutions in most if not all cases.
Where databases are used both operationally (including for master data management) and for a data warehouse, the same considerations may apply – even though we are now in cases where the types of operation (e.g., updates vs. querying) and the types of data (e.g., customer vs. financial) may be somewhat different. One-way replication with its attendant ETL-style data cleansing is only one way to coordinate the overall performance of multiple instances, not to mention queries spanning them.  DV’s added flexibility gives users the ability to optimize better in many cases across the entire set of use cases.
Again, this advantage may not have been perceived (a) because not many implementers are focused on the multiple-copy case and (b) because DV implementation is probably compared against performance against each individual instance instead of or as well as against the multiple-instance database as a whole. Nevertheless, at least theoretically, this performance advantage should appear – and especially because, in this case, the “extra software layer” should not typically add some DV performance cost.

The User Bottom Line:  Where’s The Pain?


It seems that, theoretically at least, we might expect to see actual performance gains over the next 1-2 years over “business as usual” from DV implementation in the majority of use cases, and that this proportion should increase, both over time after implementation and as corporations’ information architectures continue to elaborate. The key to detecting these advantages now, if the IT shop is willing to do it, is more sophisticated metrics about just what constitutes a performance hit or a performance improvement, as described above.
So maybe there isn’t such a performance tradeoff for all the undoubted benefits of DV, after all.  Or maybe there is.  After all, there is a direct analogy here with agile software development, which seems to lose by traditional metrics of cost efficiency and quality attention, and yet winds up lower-cost and higher-quality after all.  The key secret ingredient in both is ability to react or proact rapidly in response to a changing environment, and better metrics reveal that overall advantage. But the “tradeoff” for both DV and agile practices may well be the pain of embracing change instead of reducing risk.  Except that practitioners of agile software development report that embracing change is actually a lot more fun.  Could it be that DV offers a kind of support for organizational “information agility” that has the same eventual effect: gain without pain?
Impossible. Gain without pain. Perish the very thought.  How will we know we are being organizationally virtuous without the pains that accompany that virtue?  How could we possibly improve without sacrifice? 
Well, I don’t know the answer to that one.  However, I do suggest that maybe, rather than the onus being on DV to prove it won’t torch performance, the onus should perhaps be on those advocating the status quo, to prove it will.  Because it seems to me that there are plausible reasons to anticipate improvements, not decreases, in real-world performance from DV.

Wednesday, October 10, 2012

Data Virtualization: Users Grok What I Said 10 Years Ago


Ten years ago I put out the first EII (now data virtualization) report.  In it I said:
  • ·         The value of DV is both in its being a database veneer across disparate databases, and in its discovery and storage of enterprise-wide global metadata
  • ·         DV can be used for querying and updates
  • ·         DV is of potential value to users and developers and administrators. Users see a far wider array of data, and new data sources are added more quickly/semi-automatically. Developers have “one database API to choke.” Administrators can use it to manage multiple databases/data stores at a time, for cost savings
  • ·         DV can be great for mergers, metadata standardization, and as a complement to existing enterprise information architectures (obviously, including data warehousing)
  • ·         DV is a “Swiss army knife” that can be used in any database-related IT project
  • ·         DV is strategic, with effects not just on IT costs but also on corporate flexibility and speed to implement new products (I didn’t have the concept of business agility then)
  • ·         DV can give you better access to information outside the business.
  •  DV can serve as the “glue” of an enterprise information architecture.
I’m at DV Day 2012 (Composite Software) in NYC. Today, for the first time, I have heard not just vendors talking about implementing these things, but users actually doing them – global metadata, user self-service access to the full range of corporate data, development using SQL access, use for updating in operational databases, use for administration of at least the metadata management of multiple databases as well as archiving administrative tasks, use in mergers, use in “data standardization”, use as an equal partner with data warehousing, use in just about every database-related IT project, selling by as strategic to the business with citations of impacts on the bottom line through strategic products as well as “business agility”, use to access extra-enterprise Big Data across multiple clouds, and use as the framework of an enterprise information architecture.

I just wanted to say, on a personal note, that this is what makes being an analyst worthwhile.  Not that I was right 10 years ago. That, for once, by the efforts of everyone pulling together to make the point and elaborate the details, we have managed to get the world to implement a wonderful idea. I firmly believed that, at least in some small way, the world would be better off if we implemented DV. Now, it’s clear that that’s true, and that DV is unstoppable.

And so, I just wanted to take a moment to celebrate.

Sunday, September 23, 2012

Is Your Organization Suffering From Data Warehouse Disease?


I have a feeling that a fair amount of readers – especially vendors and IT BI types – are going to be upset by what I have to say in this post.  However, viewing some of the material that has passed across my desk recently, I really think it’s time to raise the question of whether too much organizational power given to data warehouse folks is beginning to cause some significant under-performance in meeting today’s key organizational information management needs.

The immediate occasion for these reflections is that I am partway through a book on a related subject that goes into some detail on data warehousing’s view of the world:  how BI should be handled, what the organizational information architecture should be, and how we got this way.  This book will remain nameless, because in many ways it’s an excellent primer.  However, over the last 22-31 years (depending on whether you count my software development days), I have had a cross-organization, cross-vendor view of the same area, and I have to say that the book redefines history and the purposes of various things in the ideal information architecture in major ways.

Usually, I find that going over history just wastes time in a blog post – but here, it helps to see how data warehouse concepts of common information management terms make them reinterpret the purposes of the underlying products, making the information architecture – and the whole information handling process – potentially (and, probably, actually) less effective in the medium and long term. So let’s combine history and exposition of my assertion.

A Data Warehousing View of the World

In brief, the book’s view of the information architecture seems to be as follows: Data of all types comes in to production systems, which immediately pass it on to the data warehouse for cleansing and aggregation. Behind the data warehouse is an optional operational data store for key data, and things like master data management operate in parallel with the data warehouse to provide a global view of multiple local ways to store customer data. On top of the data warehouse are key Business Intelligence applications, which include both repetitive, scheduled reporting and analytics.

Now, this view of the world seems reasonable if you were born yesterday, or if you’ve spent the last fifteen years entirely in data warehousing.  However, there are, in my view, some major problems with it.
In the first place, afaik, only in data warehousing are the databases at the initial entry point referred to as “production systems”.  For twenty years, I have been calling them “operational databases”. In fact, they were business-critical before data warehousing existed, and so were the apps on top of them – like ERP. 

Why does this matter? Because it allows data warehouse folks to shift the “operational data store” behind the data warehouse.  The operational data store is a later concept, and one that I (among others, I assume) wrote papers proposing around 2004 and 2005. The idea is that the data warehouse is simply too slow to react immediately to key operational data – but that operational data is scattered across multiple operational data stores, and so an “operational data store” makes sure that a subset of operational data for quick decision-making is either put in a central point for quick analysis in parallel with its arrival, or monitored by a central “virtual database.” Putting the operational data store behind the data warehouse defeats its entire purpose.

Likewise, the master data management system. I wrote papers on this in assessing IBM’s version of the concept in 2006 and 2007. Again, the notion was of combining operational data coming in to operational databases – in this case, by enforcing a common format that allowed cross-organization and cross-country leveraging of operational data by ERP and customer intelligence apps. By redefining the master data management as existing within the data warehouse or at the same remove from operational databases, data warehouse folks ensure that master data management moves no faster than the data warehouse.

And finally, there is the idea that (implicitly) analytics is entirely contained in BI, and hence is entirely dependent on the data warehouse. On the contrary, an increasing amount of analytics goes on outside of BI.  For example, analytics is part of products that analyze computer infrastructure semi-automatically to optimize performance or detect upcoming problems. Or, it is used to analyze key computer-supported business processes.  This is “intelligence” in the sense of “military intelligence” – proactively going out and finding out what’s going on – but it is not “business intelligence” in the sense of finding out what’s going on inside and outside the business on the basis of data that is handed to you, and that your reporting tools are too slow or shallow to tell you. In other words, these applications of analytics are entirely outside of a reactive data warehouse.

Why It Matters

There are two places that over-emphasis on data warehousing can impede organizational BI and other information management effectiveness:  the information architecture, and the organization’s “agility” in responding to new kinds of information from outside. As I’ve suggested in the previous section, a data warehousing view of the information architecture shifts operations that involve lots of “updates” and data just arrived from outside to the data warehouse or behind it.  That means going through the data-warehouse cleansing and aggregation process and arriving in a centralized location that is handling queries from all over the organization and is optimized for adding new data not “on the fly” but in delayed bursts. There is simply no way that is going to be as timely as performing tasks on the data as it arrives in the operational systems.

Just as troubling, the entire emphasis of the organization is now more reactive and focused farther away from the organization’s “antennae” to the outside environment. The IT organization appears to be focused on responding to new demands from business for timelier data, not actively seeking the latest new information and merging it back into existing systems. The IT organization appears to emphasize cleaning up the data and merging it and only then analyzing it at an internal “choke point”, rather than handling the information faster where it arrives. 

If you think these concerns are theoretical, think about the case of social-media Big Data. Yes, Oracle as a major vendor is emphasizing inhaling huge amounts of this data from multiple clouds into the data warehouse and then analyzing it – when the whole purpose of the NoSQL movement is to allow rapid in-cloud analysis of inconsistent, uncleansed data – but it would not do so unless there was some organizational push to avoid analytics outside the data warehouse.  I conclude that there is some strong evidence that a data warehousing focus is impeding organizational ability to process and feed to business decision makers key information in as timely a fashion as possible.

Moreover, there is some sense that this is not an organizational quirk but a tendency so embedded in the IT organization that this impediment is a symptom not of a temporary problem that is easy to fix, but rather of an organizational “disease.” In other words, simply directing the organization to pay more attention to doing social-media processing in the cloud will probably not work.

Action Strategies and Conclusion

First (although I think there is little danger of this) I must caution against throwing the baby out with the bathwater.  There are very good reasons to have a data warehouse performing the core functions of querying for BI. I have, in the past, conjectured that if I were to design a new information architecture today, I might not create a data warehouse or data mart at all – instead, I might impose “data virtualization” and master data management tools over existing operational databases. However, practically speaking, in most if not all cases, the sheer experience behind today’s data warehousing products makes them far more preferable for core functions.

Rather, I would suggest that data warehousing be placed under, and be responsive to rather than dominant over, an information architecture and information strategy function aimed more at the edge of the organization than its central data center. This is not a matter of making the organization more responsive to the business; it is a matter of making the IT organization more agile (by my definition, which stresses the utility of proactive and outside-the-organization-directed agility).

Until I saw this book, which suggested that data warehouse folks had gone too far in asserting “IT information handling is all about the data warehouse”, I was not too concerned about data warehousing folks; I would get into annoying arguments with folks who thought I just didn’t “get” data warehousing, but it seemed to me that the benefits of a powerful database-related IT function outweighed the negatives of data warehouse folks’ “not invented here” blind spots.  Now, I am rethinking my position.  If the result of this type of rewriting of history is an increasingly sub-optimal information architecture, then such a “disease” is not so harmless after all.

Does your organization suffer from data warehouse disease?  If so, what do you think should be done about it?

Wednesday, March 28, 2012

Oracle, EMC, IBM, and Big Data: Avoiding The One-Legged Marathon

Note: this is a post of an article published in Nov. 2011 in another venue, and the vendors profiled here have all upgraded their Big Data stories significantly since then. Imho, it remains useful as a starting point for assessing their Big Data strategies, and deciding how to implement a long-term IT Big Data strategy oneself.

In recent weeks, Oracle, EMC, and IBM have issued announcements that begin to flesh out in solutions their vision of Big Data, its opportunities, and its best practices. Each of these vendor solutions has significant advantages for particular users, and all three are works in progress.

However, as presented so far, Oracle’s and EMC’s approaches appear to have significant limitations compared to IBM’s. If the pair continues to follow the same Big Data strategies, I believe that many of their customers will find themselves significantly hampered in dealing with certain types of Big Data analysis over the long run – an experience somewhat like voluntarily restricting yourself to one leg while running a marathon.

Let’s start by reviewing the promise and architectural details of Big Data, then take a look at each vendor’s strategy in turn.

Big Data, Big Challenges
As I noted in an earlier piece, some of Big Data is a relabeling of the incessant scaling of existing corporate queries and their extension to internal semi-structured (e.g., corporate documents) and unstructured (e.g., video, audio, graphics) data. The part that matters the most to today’s enterprise, however, is the typically unstructured data that is an integral part of customers’ social-media channel – including Facebook, instant messaging, Twitter, and blogging. This is global, enormous, very fast-growing, and increasingly integral (according to IBM’s recent CMO survey) to the vital corporate task of engaging with a "customer of one" throughout a long-term relationship.

However, technically, handling this kind of Big Data is very different from handling a traditional data warehouse. Access mechanisms such as Hadoop/MapReduce combine open-source software, large amounts of small or PC-type servers, and a loosening of consistency constraints on the distributed transactions (an approach called eventual consistency). The basic idea is to apply Big Data analytics to queries where it doesn’t matter if some users get "old" rather than the latest data, or if some users get an answer while others don’t. As a practical matter, this type of analytics is also prone to unexpected unavailability of data sources.

The enterprise cannot treat this data as just another BI data source. It differs fundamentally in that the enterprise can be far less sure that the data is current – or even available at all times. So, scheduled reporting or business-critical computing based on Big Data is much more difficult to pull off. On the other hand, this is data that would otherwise be unavailable for BI or analytics processes – and because of the approach to building solutions, should be exceptionally low-cost to access.

However, pointing the raw data at existing BI tools would be like pointing a fire hose at your mouth, with similarly painful results. Instead, the savvy IT organization will have plans in place to filter Big Data before it begins to access it.

Filtering is not the only difficulty. For many or most organizations, Big Data is of such size that simply moving it from its place in the cloud into an internal data store can take far longer than the mass downloads that traditionally lock up a data warehouse for hours. In many cases, it makes more sense to query on site and then pass the much smaller result set back to the end user. And as the world of cloud computing keeps evolving, the boundary between "query on site" and "download and query" keeps changing.

So how are Oracle, EMC, and IBM dealing with these challenges?

Oracle: We Control Your Vertical
Reading the press releases from Oracle OpenWorld about Oracle’s approach to Big Data reminds me a bit of the old TV show Outer Limits, which typically began with a paranoia-inducing voice intoning "We control your horizontal, we control your vertical …" as the screen began flickering.

Oracle’s announcements included specific mechanisms for mass downloads to Oracle Database data stores in Oracle appliances (Oracle Loader for Hadoop), so that Oracle Database could query the data side-by-side with existing enterprise data, complete with data-warehouse data-cleansing mechanisms.

The focus is on Oracle Exalytics BI Machine, which combines Oracle Database 11g and Oracle’s TimesTen in-memory database for additional BI scalability. In addition, there is a "NoSQL" database that claims to provide "bounded latency" (i.e., it limits the "eventual" in "eventual consistency"), although how it should combine with Oracle’s appliances was not clearly stated.

The obvious advantage of this approach is integration, which should deliver additional scalability on top of Oracle Database’s already-strong scalability. Whether that will be enough to make the huge leap to handling hundred-petabyte data stores that change frequently remains to be seen.

At the same time, these announcements implicitly suggest that Big Data should be downloaded to Oracle databases in the enterprise, or users should access Big Data via Oracle databases running in the cloud, but provide no apparent way to link cloud and enterprise data stores or BI. To put it another way, Oracle is presenting a vision of Big Data used by Oracle apps, accessed by Oracle databases using Oracle infrastructure software and running on Oracle hardware with no third party needed. We control your vertical, indeed.

What also concerns me about the company’s approach is that there is no obvious mecha-nism either for dealing with the lateness/unavailability/excessive lack of quality of Big Data, or for choosing the optimal mix of cloud and in-enterprise data location. There are only hints: the bounded latency of Oracle NoSQL Database, or the claim that Oracle Data Integrator with Application Adapter for Hadoop can combine Big Data with Oracle Database data – in Oracle Database format and in Oracle Database data stores. We control your horizontal, too. But how well are we controlling it? We’ll get back to you on that.

EMC: Competing with the Big Data Boys
The recent EMC Forum in Boston in many ways delivered what I regard as very good news for the company and its customers. In the case of Big Data, its acquisition of Greenplum with its BI capabilities led the way. And Greenplum finally appears to have provided EMC with the data management smarts it has always needed to be a credible global information management solutions vendor. In particular, Greenplum is apparently placing analytical intelligence in EMC hardware and software, giving EMC a great boost in areas such as being able to handle querying within the storage device and monitoring distributed systems (such as VCE’s VBlocks) for administrative purposes. These are clearly leading-edge, valuable features.

EMC’s Greenplum showed itself to be a savvy supporter of Big Data. It supports the usual third-party suspects for BI: SQL, MapReduce, and SAS among others. Querying is "software shared-nothing", running in virtual machines on commodity VBlocks and other scale-out/grid x86 hardware. Greenplum has focused on the fast-deploy necessities of the cloud, claiming a 25-minute data-model change – something that has certainly proved difficult in large-scale data warehouses in the past.

Like Oracle, Greenplum offers "mixed" columnar and row-based relational technology; un-like Oracle, it is tuned automatically, rather than leaving it up to the customer how to mix and match the two. However, its answer for combining Big Data and enterprise data is also tactically similar to Oracle’s: download into the Greenplum data store.

Of our three vendors, EMC via Greenplum has been the most concrete about what one can do with Big Data, offering specific support for combining social graphing, customer-of-one tracking of Twitter/Facebook posts, and mining of enterprise customer data. The actual demo had an unfortunate "1984" flavor, however, with a customer’s casual chat about fast cars being used to help justify doubling his car insurance rate.

The bottom line with Greenplum appears to be that its ability to scale is impressive, even with Big Data included, and it is highly likely to provide benefits out of the box to the savvy social-media analytics implementer. Still, it avoids, rather than solves, the problems of massive querying across Big Data and relational technology – it assumes massive downloads are possible and data is "low latency" and "clean", where in many cases it appears that this will not be so. EMC Greenplum is not as "one-vendor" a solution as Oracle’s, but it does not have the scalability and robustness track record of Oracle, either.

IBM: Avoiding Architectural Lock-In
At first glance, IBM appears to have a range of solutions for Big Data similar to Oracle and EMC – but more of them. Thus, it has the Netezza appliance; it has InfoSphere BigInsights for querying against Hadoop; it has the ability to download data into both its Informix/in-memory technology and DB2 databases for in-enterprise data warehousing; and it offers various Smart Analytics System solutions as central BI facilities.

Along with these, it provides Master Data Management (InfoSphere MDM), data-quality features (InfoSphere DataStage), InfoSphere Streams for querying against streaming Web sensor data (like mobile GPS), and the usual quick-deployment models and packaged hardware/software solutions on its scale-out and scale-up platforms. And, of course, everyone has heard of Watson – although, intriguingly, its use cases are not yet clearly pervasive in Big Data implementations.

To my mind, however, the most significant difference in IBM’s approach to Big Data is that it offers explicit support for a wide range of ways to combine Big-Data-in-place and enterprise data in queries. For example, IBM’s MDM solution allows multiple locations for customer data and supports synchronization and replication of linked data. Effectively used, this allows users to run alerts against Big Data in remote clouds or dynamically shift customer querying between private and public clouds to maximize performance. And, of course, the remote facility need not involve an IBM database, because of the MDM solution’s cross-vendor "data virtualization" capabilities.

Even IBM’s traditional data warehousing solutions are joining the fun. The IBM IOD conference introduced the idea of a "logical warehouse", which departs from the idea of a single system or cluster that contains the enterprise’s "one version of the truth", and towards the idea of a "truth veneer" that looks like a data warehouse from the point of view of the analytics engine but is actually multiple operational, data-warehouse, and cloud data stores. And, of course, IBM’s Smart Analytics Systems run on System x (x86), Power (RISC) and System z (mainframe) hardware.

On the other hand, there are no clear IBM guidelines for optimizing an architecture that combines traditional enterprise BI with Big Data. It gives one the strong impression that IBM is providing customers with a wide range of solutions, but little guidance as to how to use them. That IBM does not move the customer towards an architecture that may prevent effective handling of certain types of Big-Data queries is good news; that IBM does not yet convey clearly how these queries should be handled, not so much.

Composite Software and the Missing Big-Data Link
One complementary technology that users might consider is data virtualization (DV), as provided by vendors such as Composite Software and Denodo. In these solutions, subqueries on data of disparate types (such as Big Data and traditional) are optimized flexibly and dynamically, with due attention to "dirty" or unavailable data. DV solution vendors’ accrued wisdom can be simply summed up: usually, querying on-site instead of doing a mass download is better.

How to deal with the "temporal gap" between "eventually consistent" Big Data and "hot off the press" enterprise sales data is a matter for customer experimentation and fine-tuning, but that customer can always decide what to do with full assurance that subqueries have been optimized and data cleansed in the right way.

The Big Data Bottom Line
To me, the bottom line of all of this Big Data hype is that there is indeed immediate business value in there, and specifically in being able to go beyond the immediate customer interaction to understand the customer as a whole and over time – and thereby to establish truly win-win long-term customer relationships. Simply by looking at the social habits of key consumer "ultimate customers," as the Oracle, EMC, and IBM Big Data tools already allow you to do, enterprises of all sizes can fine-tune their interactions with the immediate customer (B2B or B2C) to be far more cost-effective.

However, with such powerful analytical insights, it is exceptionally easy to shoot oneself in the foot. Even skipping quickly over recent reports, I can see anecdotes of "data overload" that paralyze employees and projects, trigger-happy real-time inventory management that actually increases costs, unintentional breaches of privacy regulations, punitive use of newly public consumer behavior that damages the enterprise’s brand or perceived "character", and "information overload" on the part of the corporate strategist.

The common theme running through these user stories is a lack of vendor-supplied context that would allow the enterprise to understand how to use the masses of new Big Data properly, and especially an understanding of the limitations of the new data.

Thus, in the long run, the best performers will seek a Big-Data analytics architecture that is flexible, handles the limitations of Big Data as the enterprise needs them handled, and allows a highly scalable combination of Big Data and traditional data-warehouse data. So far, Oracle and EMC seem to be urging customers somewhat in the wrong direction, while IBM is providing a "have it your way" solution; but all of their solutions could benefit strongly from giving better optimization and analysis guidance to IT.

In the short run, users will do fine with a Big-Data architecture that does not provide an infrastructure support "leg" for a use case that they do not need to consider. In the long run, the lack of that "leg" may be a crippling handicap in the analytics marathon. IT buyers should carefully consider the long-run architectural plans of each vendor as they develop.

Wednesday, January 25, 2012

The Other BI: Oracle TimesTen and In-Memory-Database Streaming BI

This blog post highlights a software company and technology that I view as potentially useful to organizations investing in business intelligence (BI) and analytics in the next few years. Note that, in my opinion, this company and solution are not typically “top of the mind” when we talk about BI today.

The Importance of TimesTen-Type In-Memory Database Technology to BI

All right, now I’m really stretching the definition of “other”. Let’s face it, Oracle is “top of the mind” when we talk about BI, and they recently announced a TimesTen appliance, so TimesTen is not an invisible product, either. And finally, the hoopla about SAP HANA means that in-memory database technology itself is probably presently pretty close to the center of IT’s radar screen.

So why do I think Oracle’s TimesTen is in some sense not “top of the mind”? Answer: because there are potential applications of in-memory databases in BI for which the technology itself, much less any vendor’s in-memory database solution, is not a visible presence. In particular, I am talking about in-memory streaming databases.

To understand the relevance of in-memory databases to complex event processing and BI, let’s review the present use cases of in-memory databases. Originally, in-memory technology was just the thing for analyzing medium-scale amounts of financial-market information in real time, information such as constantly changing stock prices. Lately, in-memory databases have added two more BI duties: (a) serving as a “cache” database for enterprise databases, to speed up massive BI where smaller chunks of data could be localized, and (b) serving as a really-high-performance platform for mission-critical small-to-medium-scale BI applications that require less scaling year-to-year, such as some SMB reporting. These new tasks have arrived because rapid growth in main-memory storage has inevitably allowed in-memory databases to tackle a greater share of existing IT data-processing needs. To put it another way, when you have an application that is always going to require 100 GB of storage, sooner or later it makes sense to use an in-memory database and drop the old disk-based one, because in-memory database performance will typically be up to 10-100 times faster.

Now let’s consider event-processing or “streaming” databases. Their main constraint today in many cases is how much historical context they can access in real-time in order to deepen their analysis of incoming data before they have to make a routing or alerting decision. If that data can be accessed in main memory instead of disk, effectively up to 10-100 times the amount of “context” information can be brought to bear in the analysis in the same amount of time.

In other words, for streaming BI, IT potentially has two choices – a traditional event-processing database that is often entirely separate from a back-end disk-based database, or (2) a traditional main-memory database already pre-optimized for in-depth main-memory analysis and usually pre-integrated with a disk-based database (as TimesTen is with Oracle Database) as a “cache database” in cases where disk must be accessed. How to choose between the two? Well, if you don’t need much historical context for analysis, the event-processing database probably has the edge – but if you’re looking to upgrade your streaming BI, that’s not likely to be the case. In other cases, such as those where the processing is “routing-light” and “analysis-heavy”, an in-memory database not yet optimized for routing but far more optimized for in-depth analytics performance would seem to make more sense.

Thus, one way of looking at the use case of in-memory database event processing is to distinguish between in-enterprise and extra-enterprise data streams (more or less). Big Data is an example of an extra-enterprise stream, and can involve a fire hose of “sensor-driven Web” (GPS) and social media data that needs routing and alerting as much as it needs analytics. Business-critical-application-destined and embedded-analytics data streams are an example of in-enterprise data, even if admixed with a little extra-enterprise data; they require heavier-duty cross-analysis of smaller data streams. For these, the in-memory database’s deeper analysis before a split-second decision is made is probably worth its weight in gold, as it is in the traditional financial in-memory-database use case.

Won’t having two databases carrying out the general task of handling streaming data complicate the enterprise architecture? Not really. Past experience shows us that using multiple databases for finer-grained performance optimization actually decreases administrative costs, since the second database, at least, is typically much more “near-lights-out,” while switching between databases doesn’t affect users at all, because a database is infrastructure software that presents the same standard SQL-derivative interfaces no matter what the variant. And, of course, the boundary between event-processing database use cases and in-memory ones is flexible, allowing new ways of evolving performance optimization as user needs change.

The Relevance of Oracle TimesTen to Streaming BI

In many ways, TimesTen is the granddaddy of in-memory databases, a solution that I have been following for fifteen years. It therefore has leadership status in in-memory database use-case experience, and especially in the financial-industry stock-market-data applications that resemble my streaming-BI use case as described above. What Oracle has added since the acquisition is database-cache implementation and experience, especially integrated with Oracle Database. At the same time, TimesTen remains separable at need from other Oracle database products, as in the new TimesTen Appliance.

These characteristics make TimesTen a prime contender for the potential in-memory streaming BI market. Where SAP HANA is a work in progress, and approaches like Volt are perhaps less well integrated with enterprise databases, TimeTen and IBM’s solidDB stand out as combining both in-memory original design and database-cache experience – and of these two, TimesTen has the longer in-memory-database pedigree.

It may seem odd of me to say nice things about Oracle TimesTen, after recent events have raised questions in my mind about Oracle BI pricing, long-term hardware growth path, and possible over-reliance on appliances. However, inherently an in-memory database is much less expensive than an enterprise database. Thus, users appear to have full flexibility to use TimesTen separately from other Oracle solutions, free from worries about possible long-term effects of vendor lock-in.

Potential Uses of TimesTen-Type In-Memory Streaming BI for IT

As noted above, the obvious IT use cases for TimesTen-type streaming BI lie in driving deeper analysis in in-enterprise streaming applications. In particular, in the embedded-analytics area, in-memory performance speedups can allow consideration of a wider array of systems-management data in fine-tuning downtime-threat and performance-slowdown detection. In the real-time analytics area, an in-memory database might be of particular use in avoiding “over-steering”, as when predictable variations in inventories cause overstocking because of lack of historical context. In the Big Data area, an in-memory database might apply where the data has been pre-winnowed to certain customers, and a deeper analysis of those customers fine-tunes an ad campaign. For example, within a half-hour of the end of the game, Dick’s Sporting Goods had sent me an offer of a Patriots’ AFC Championship T-shirt, complete with visualization of the actual T-shirt – a reasonably well-targeted email. That’s something that’s far easier to do with an in-memory database.

IT should also consider the likely evolution of both event-processing and in-memory databases over the next few years, as their capabilities will likely become more similar. Here, the point is that event-processing databases often started out not with data-management tools, but with file-management ones – making them significantly less optimized “from the ground up” for analysis of data in main memory. Still, event-processing databases such as Progress Apama may retain their event-handling, routing, and alerting advantages, and thus the situation in which in-memory is better for in-enterprise and event-processing is better for extra-enterprise is likely to continue. In the meanwhile, increasing use of in-memory databases for the older use cases cited above means that in-memory streaming-BI databases offer an excellent way of gaining experience in their use, before they become ubiquitous. That, in turn, means that narrow initial “targets of opportunity” in one of the situations cited in the previous paragraph are a good idea, whatever the scope of one’s overall in-memory database commitment right now.

The Bottom Line for IT Buyers

In some ways, this is the least urgent and most speculative of the “other BI” solutions I have discussed so far. We are, after all, discussing additional performance and deeper analytics in a particular subset of IT’s needs, and in an area where the technology of in-memory databases and their event-processor alternatives is moving ahead rapidly. In a sense, this is really an opportunity for those IT shops that specialize in applying a little extra effort and “designing smarter” across multiple new technologies to provide a nice ongoing competitive advantage. For the rest, if the shoe can easily be made to fit, why not wear it?

My suggestion for most IT buyers, therefore, is therefore to have a “back-pocket” in-memory-database-for-streaming-BI short list that can be whipped out at the appropriate time. Imho, Oracle TimesTen right now should be on that list.

I hate to close without noting the overall long-term BI potential of in-memory databases. The future of in-memory databases is not, in my firm opinion, to supersede the IBM DB2s, Oracle Databases, and Microsoft SQL Servers of the world, at any time in the next four years. The hardware technologies to enable such a thing are not yet clear, much less competitive. Rather, the value of in-memory databases is to allow us to optimize our querying for both main-memory and disk storage – which are two very different things, and which will both apply appropriately to many key customer needs over the next few years. Overall, the effect will be another major ongoing jump in data-processing performance. As we enter this new database-technology era, those who initially kick the tires in a wider variety of BI projects will find themselves with a significant “experience” advantage over the rest, especially because the key to outstanding success will be determining the appropriate boundary between disk-based and in-memory database usage. Don’t force in-memory streaming BI into the organization. Do keep checking to see if it will fit your immediate needs. Sooner or later, it probably will.

Friday, January 13, 2012

The Other BI: HP Vertica and Columnar Databases

This blog post highlights a software company and technology that I view as potentially useful to organizations investing in business intelligence (BI) and analytics in the next few years. Note that, in my opinion, this company and solution are not typically “top of the mind” when we talk about BI today.

The Importance of Vertica-Type Columnar Database Technology to BI

Last year, I wrote a blog post saying that it was likely that HP would underestimate the columnar database technology in Vertica, and if so they were missing a major opportunity. In the last year, HP has been pretty quiet about Vertica, but I have partially changed my mind, to the point where I want to call attention to Vertica as a less visible candidate for IT buyers to get the full benefits of columnar database technology over the next 2-3 years.

Let’s start with columnar technology. Here, I want to go more in-depth into Vertica’s core technology than usual, because it’s an excellent way to begin to see the benefits of columnar beyond traditional row-oriented databases.

The original idea of Vertica was to recast the relational database to focus on the (data warehousing) case where there are few if any updates. The redesign started with the idea that the data should be stored in "columns" rather than rows; the details of this are that the columns themselves (because they don't have to follow relational dogma) can be stored in a highly compressed format, with lots of compression techniques like inverted list, bit-mapped indexing, and hashing, as appropriate. Thus, (a) the database can use the column format to zero in faster on the data that the query is gathering, (b) because the data is compressed an average of 10 times (according to Vertica), more data can be crammed into main memory for faster processing. Result: a claimed 10-100 times speedup in performance, comparable to in-memory databases but far more scalable. It also means the database can handle at least 10 times more data (say, 100 terabytes instead of 5) with the same performance for a given query; or that the data center can use an order of magnitude less storage.

Now, all this does not come without a cost, and the typical cost would at first seem to be speed of updating. That is, the column storage format requires more revision of the data stored on disk when an update arrives, so update is slower. But this is counteracted by the ability to load more of localized data at once into main memory in a compressed form, for faster in-memory updating. Only at update frequencies typical of old-style operational online transaction processing (OLTP) does the row-oriented relational database have a clear edge.

The elaboration of the design in Vertica is that the basic data is also stored as "projections" (aka materialized views). That is, a set of columns in a tuple is stored one (relational) way; each column also shows up in a projection, but the projection is cross-tuple (one from tuple A, one from tuple B, etc.). This accomplishes two things: one, it gives an alternative way of querying which may be faster than basic storage, and two, it gives redundancy and therefore robustness, in a similar way to RAID 5 (projections can be "striped" across disks).

Now, here's where things get really interesting. Practically speaking, today, in data-warehousing-type databases, updates via "load windows" are becoming more and more frequent, to the point where data is pretty up-to-date and updates are a bigger part of data warehousing. To keep "write locks" from gumming up performance (especially with column update being slower), Vertica splits the storage into a write-optimized column store (WOS; effectively, a cache) and a Read-optimized Column Store (ROS). Periodically, the WOS becomes the ROS. So the write locks for the updates only interfere with reads when there’s a mass update. At the same time, such a mass update can re-store whole chunks of the ROS for optimum storage efficiency. Moreover, to gain currency, the query can be carried out across the ROS and WOS. And, because there is all this redundancy, there is no need for logs—another performance improvement. Note that because of its redundancy, Vertica doesn't need to do roll-back/roll-forward nor backup/restore.

The net of all this for IT buyers is that columnar databases in general, and Vertica in particular, should be able to deliver on average much better performance than traditional relational databases in the majority of not-highly-update-intensive cases, due mostly to its compression abilities, and that addition of other technologies like in-memory technology to both alternatives will not alter this superiority.

The Relevance of HP Vertica to BI

This kind of approach cries out for integration with or development of sophisticated admin tools, expansion beyond data warehousing and analytics to “mixed” transactions in competition with the noSQL fad, better programming tools to build up a war chest of business/industry customized solutions, and using a relational database as an OLTP complement. The resulting data-management platform would be a solid alternative for all sizes of enterprise to the “relational fits all” or “let the thousand flowers bloom” strategies of most organizations.

Once this platform is in place, it needs to become the keystone of enterprise architectures, not just an analytics or business intelligence “super-scaling” engine. That means adding integration with semi-structured and unstructured data. It also means adding major functionality for handling content, and integration with storage software for additional performance optimization. And so, anticipating that HP would not do this, I criticized the HP acquisition of Vertica last year.

Well, two things happened: HP did more than I thought it would, and competitors did less. HP bought a company called Autonomy, which added semi-structured/unstructured data support. Necessarily, this takes Vertica beyond pure data-warehousing-style analytics into a more update-intensive world, and HP’s redirection of Mercury Interactive towards agile ALM (application lifecycle management) associated Vertica with better programming tools. Meanwhile, SAP took its eye off Sybase IQ with its focus on HANA, IBM at least temporarily walked away from its Netezza semi-columnar database technology, and Oracle’s columnar-optional appliance ran into questions about its long-term hardware growth path. In other words, the result of half a loaf from HP and less than half a loaf from everyone else is that Vertica is moving towards leadership status in delivering columnar database technology to all scales of BI and analytics.

Meanwhile, of course, only the deluded think that HP will suddenly vanish, while database technology and the rest of the new software embed themselves ever deeper in HP’s DNA. HP Vertica is going to be around for quite a while; and it will be an attractive option for quite a while.

Potential Uses of Vertica-Type Columnar-Based BI for IT

The use cases of a columnar database IT is straightforward. IT should use a columnar database in new projects as an alternative or complement to a traditional relational database, unless the operations are update-intensive, in which case row-oriented relational is preferred. As a complement, columnar databases operate on a “switching” basis, in which an overall engine decides which queries should be allocated to row-oriented, which to columnar, usually on the basis of whether two or more of the “fields” involved in an operation can be compressed highly by using a columnar format. Oracle (and, until recently, IBM Netezza) takes this approach; but IT can also do its own switching mechanism.

And that’s it. Over the next 2-3 years, if not already, columnar can scale as high as querying, can integrate with as many data types and upper-level tools and applications, and can evolve to greater performance/scalability just as rapidly as the traditional row-oriented database. In the long run, in a lot of use cases, and sometimes in the short run, that favors Vertica-type columnar.

However, right now, columnar requires in some cases to “grow into” its assigned role in a new project, by adding administrative tools for particular cases. Therefore, in most applications where 24x7 operation and an adequate level of customer response time is business-critical, relational row-oriented should still be preferred. That should leave plenty of analytical and other BI uses for which Vertica-type columnar database software will deliver an important performance advantage.

The Bottom Line for IT Buyers

Over the next few years, IT buyers can take one of two views: the author of this blog post is prescient, columnar will replace row-oriented in the majority of new applications in BI and other areas, and we should include columnar in all our short lists from now on; or, the author of this blog post is wrong about the future, but columnar is useful for some things right now, and trying to standardize on one database is a fool’s game that we no longer bother to try to play. If IT buyers hold the second view, then they should be focused on applying columnar to analysis of huge amounts of structured data with “sparse” fields where high compression is achievable – like five-field customer names (Mr. John Taylor Jakes, Jr.) and product codes. Spend the resulting improvements on increased performance, lowered storage costs, or both.

Again, this is not a matter of a pre-short list, unless you have a “gray area” BI project involving somewhat update-intensive or somewhat business-critical little-downtime apps, in which case you want to wait for columnar to evolve a little. In all other cases, HP Vertica should go on the short list along with the obvious others, like Sybase IQ. Right now, Vertica appears to be ahead both in some of the needed features to adapt to new analytics needs and in speed of evolution. One never knows – but over the next year, that leadership role may continue.

Above all, IT buyers should not listen to any FUD from traditional relational vendors suggesting that this is yet another new technology, like object databases, that will eventually fall to earth with a thud. Columnar database technology proved its superiority in many situations long ago in the non-relational world, with CCA’s Model 204, and has found uses continuously since then, like bit-mapped indexing. Most times there’s a fair BI matchup, as with some of the TPC benchmarks of the last seven years, columnar comes out well ahead. Under whatever name, columnar database technology is not going away. Therefore, its markets will continue to grow relative to row-oriented relational. For IT buyers, acquiring columnar BI solutions like HP’s Vertica is simply being smart and getting a little ahead of the curve.

Thursday, January 5, 2012

The Other BI: EMC Greenplum and Embedded Analytics

This blog post highlights a software company and technology that I view as potentially useful to organizations investing in business intelligence (BI) and analytics in the next few years. Note that, in my opinion, this company and solution are not typically “top of the mind” when we talk about BI today.

The Importance of the Greenplum Software Technology to BI

I am stretching a point when I say that EMC’s Greenplum is not “top of the mind” today. EMC has done an extensive and effective job of marketing Greenplum’s virtues in dealing with Big Data. However, what I am talking about here is embedded analytics – and there, neither Greenplum nor any other vendor solution is “top of the mind” with IT today.

More specifically, I am talking about middle-tier analytics, the area that most embedded analytics will aim for in the next three years. This is not massive-data-store, in-depth-analytics BI like the data warehouse; nor is it the “smart sensor,” small-form-factor analytics that will increasingly come to the fore with the arrival of the sensor-driven Web (e.g., analytics on your iPhone). No, I am talking about medium-sized data stores, moderately in-depth analytics, and in-enterprise BI applied at the level of the department, local office, loosely-coupled storage array, or server network. This analytics does best when it is embedded in other software or in firmware, and operates semi-automatically to pick up business-process flows and alert the business before they get out of whack, or offloads load balancing from a central server. Unlike systems management software, embedded analytics not only monitors and “fixes” but also analyzes what is going on, and reports this analysis either to the top-tier data warehouse or a specific set of software, end users, and/or administrators.

Up to now, the fledgling beginnings of embedded analytics have begun to show up in the systems management software of folks like CA; but they are not separable pieces usable by other distributed software. Increasingly, the major vendors like IBM are now talking about taking analytic software from BI and analytics software suites and applying it to organization operations across the board.

However, these often involve databases retrofitted to BI in general and decision support in particular. What Greenplum represents is the obvious next step: applying a database designed from the ground up and optimized for querying and analytics. The point is that these will inevitably be better suited than data management approaches intended to handle updates as well as queries and result massaging.

This is not to say that an embedded analytics database is the end point of embedded-analytics evolution. Because most if not all available analytics databases were designed for the top tier, they are too “heavyweight” for their intended purpose: they perform more slowly, because they are tuned for much higher data-store sizes. However, whether the next turn of the market crowns a slimmed-down top-tier database or a new ground-up-designed middle-tier analytics database as the winner, either one will really do.

Over the next 2-3 years, it is reasonable for IT buyers to expect some of this technology to arrive on their doorsteps embedded in upgrades of existing solutions – but far from all of it. At some point in this period, separable analytics solutions will show up that will allow the user to go far beyond what a particular vendor is offering – if, of course, IT wants to.

Why would IT want to do this? Answer: to handle areas in which one-size-fits-all vendors are simply not moving fast enough. Take, for example, carbon accounting. Vendors have been very proactive in this area, but some of the market is moving faster still, towards monitoring that picks up on and alerts to excess emissions as they happen, and connects with the carbon accounting software when necessary. Likewise, as health care providers grapple with government mandates and Electronic Health Records, they can see coming a day in which they will need to perform damage control on breaches of privacy; but today’s tools are much slower than they could be to detect such a problem. In either case, customizable middle-tier embedded analytics that goes beyond most likely vendor offerings is needed.

The primary organization benefit of this technology, therefore, is deeper real-time understanding of in-enterprise problems that leads to better decision-making –a very cost-effective application of analytics’ general ability to improve gross margins. Embedded analytics via an analytics-adapted database may take longer to arrive than most of the Other BI that I talk about, but its advent and benefits are just as sure.

The Relevance of EMC to BI

While EMC has continued its tradition of “hands off the technology, add our markets” in the Greenplum acquisition, it has also continued another tradition: adding the technology where appropriate to its core storage software/firmware. That is, according to EMC (and I see no reason to doubt them), Greenplum technology is being put in storage controllers to offload querying from the server to the storage array. Obviously, that has a major positive implication for storage and large-BI performance. Less appreciated is the fact that this embedding of Greenplum requires that it “slim down” into a form that can operate not only on storage but also, in a middle-tier fashion, on loosely-coupled LANs serving local offices, departments, and so on. In other words, embedding on storage should mean that embedding on all other middle-tier form factors is within reach. And the acquisition of Greenplum also should mean that EMC is finally beginning to add database and BI smarts to its DNA, ensuring reasonable long-term service and support for its embedded-analytics solutions.

EMC’s market strength and apparent relative freedom from threat in the scale-out market mean that in the 2-3 year time frame I am talking about, and probably in the medium term as well, Greenplum is in no danger of going away. No, the real question for IT buyers of embedded analytics is whether EMC will have Greenplum take the next step, abstracting its slimmed-down form for embedded analytics on all vendor platforms. I can offer no guarantees of this, since it is not apparent that EMC has done such a thing before. All I can say is, if they do so, at least some sort of market will be there.

Potential Uses of Greenplum-Type Analytics for IT

It is time to point out that embedded-analytics technology is unusual in that vendors have relative freedom to delay delivering, say, multivendor or open-source middle-tier analytical databases, since it’s not high on IT wish lists. It could happen next week, or it could happen 3 years from now. So any IT acquisition of, and use of, this kind of embedded analytics will just have to wait until the vendors get around to it.

At that point, the obvious application is per-project – improving a specific business process or case-management implementation. More than other technologies, embedded analytics does not require full, integrated organizational implementation to be maximally effective. Rather, it does just fine applied to a task, a process, a function, a locality, or a local or strategic initiative. IT simply looks down the list of mission-critical projects and picks the one that benefits most from risk management or analytical automation.

The critical success factor in such projects is rapid implementation and upgrade, caused by automation of the implementation/upgrade process, allowing strategic projects a head start. Right now, while most vendors do well at this, high-end vendors like EMC seem to be setting the pace. And so, choosing EMC Greenplum (assuming it fits) in all likelihood means a better chance of rapid implementation and a database better fitted to a broad range of embedded-analytics tasks – not to mention better ongoing support for tricky cases.

The Bottom Line for IT Buyers

The IT buyer should view embedded analytics as a technology that may take a while to materialize. However, when it does, Greenplum-type embedded analytics will deliver analytics-type benefits at least equal to the whiz-bang high-end analytics now being sold – although those benefits will arrive in smaller per-project chunks. And that, in turn, means that this technology is definitely worth the IT buyer’s ongoing attention.

More specifically, the IT buyer might consider a “pre-pre-short-list” type of approach. That would involve identifying solutions such as EMC Greenplum that may wind up as part of the embedded-analytics short list in the next 2 years, and steadily moving those products in the pre-pre list over to the “pre-short list” as their technology reaches the point of usefulness (that is, it can be applied by IT rather than being embedded in another vendor solution, and it’s optimized for middle-tier analytics). Today, I would say that it appears Greenplum is probably among the closest to that take-off point. So, put it on the pre-pre short list, and get ready to put it on the short list. If everything goes right, and your CEO hits you with an urgent requirement that really demands embedded analytics, you will definitely be glad you had EMC’s Greenplum embedded analytics solution in your back pocket.

Tuesday, January 3, 2012

The Other BI: Progress Apama and Event Processing

This blog post highlights a software company and technology that I view as potentially useful to organizations investing in business intelligence (BI) and analytics in the next few years. Note that, in my opinion, this company and solution are not typically “top of the mind” when we talk about BI today.

The Importance of the Apama Software Technology to BI

The value-add of Apama to BI, in my opinion, is the value-add of applying analytics to “data in motion” on a very broad range of data. Apama carries out “event processing”: conceptually, I think of event processing as a “processing head” monitoring an Enterprise Service Bus (ESB). Most data entering the organization, as well as data moving between data stores and between users within the organization, is wrapped up as messages and sent across the ESB to its destination. In the process, the “processing head” monitors the whole stream of data-in-motion and performs analytics, alerting, and other processing based on the type of data being reviewed (or, in the aggregate, the “pattern” of a stream of related data). What is unprecedented about this kind of data processing is that (a) it focuses on data across organizational units, unlike the typical data warehouse or operational database; (b) it intercepts some data the moment it arrives in the organization, which is the ultimate in real-time business-critical data processing; and (c) it can draw a direct line between that data and a decision-maker by alerting, so that business-critical decisions can be made as quickly as possible.

Practically, of course, an “event processor” can do (c) only for a certain small subset of the information in the organization, because by itself an event-processing database does not scale nearly as much as a data warehouse. The event processor has much less historical “context” as it processes each datum, because it simply does not have the time to perform a “query from hell” on terabytes of historical data before the next datum must be processed. In-depth analytics will simply have to wait, often for an hour or more. Nevertheless, this kind of instantaneous response is, in the real world, enormously valuable when fast response to the type of events that the event processor detects from the data is indeed mission-critical and/or business-critical.

At the same time, (a) – the ability to correlate data across organizational units – is an often-underestimated value of the event processor. As the discipline of systems analysis understands, a collection of business units is as much a set of process flows between units as a set of stand-alone companies. The job of corporate is often to ensure that these process flows work well, and the value-add of the event processor in this case is to provide enterprise performance management (EPM) that reads the tea leaves of particular process flows and ties them back to glitches in the performance of the units and their coordination. In plain English, a good event processor goes beyond what you could do before because it lets you respond immediately to some new threats and opportunities in your environment, and because it tells you some of the things that are really happening to muck up your business’ overall performance as it coordinates business units. If you combine the two, you get the famous “360-degree view” inside and outside the organization.

Apama, like other event processors, never operates in a vacuum. All organizations already have operational and decision-support databases supporting key applications, and event processing must adapt itself to handle what these do not. Therefore, Apama’s value-add within the limits of (a)-(c) above can vary quite widely. Always, however, if the user does a careful analysis of the most important decisions that need to be speeded up and the gaps in business performance information, the BI done by an event processor like Apama has a major impact on the organization – not on its bottom line, necessarily, but always on its business risk. The out-of-the-blue event that businesses always face becomes much less risky when an event processor manages to detect it in a timely fashion.

Right now, businesses of all sizes are still in the early stages of use of event processing – you can tell, because case studies typically trumpet particular per-project uses. Therefore, the field is wide open for approaches such as the “event-driven architecture”, in which an event processor on top of an ESB becomes the focal point at which corporate can not only monitor but also direct information flows; the “complex event processor”, in which on-the-fly analytics becomes far deeper; and “data streaming”, in which the whole notion of a data-warehouse database is upended to be a “processing head” handling querying on multiple parallel “streams” of XML-type data. These, however, will often not achieve full implementation until 2-4 years from now, at the least. The key value-add of event processing in real-world BI over the next 2-3 years, I believe, will be the ongoing identification and implementation of the most important alerts, decisions, and business-unit correlations that it can handle, and their integration with the existing BI architecture.

The Relevance of Progress Software to BI
The relevance of Progress Software itself to BI, and to data processing in general, is less clear than in its “glory days,” when (imho) it was a pioneer in near-lights-out database administration, rapid application development, Software as a Service (SaaS), and the ESB. In those days, it had a gift for identifying the innovative infrastructure-software simplification that led its SMB/departmental customers to rapid implementation of the latest large-enterprise functionality, and beyond. Apama arrived at the end of that period, as part of a successful Progress effort to provide services and software to allow its departmental customers to scale their SMB-type technology to the division, the line of business, and even the “edge” in the data center. Apama first found its niche in rapid analysis of massive streaming financial-market data; Progress Software appears, with its Event Manager, Event Modeler process-design end-user tool, SmartBlocks, and Dashboard Studio, to have added Progress’ own strengths in SMB-driven simplicity of use. To put it another way: Apama was born large-enterprise-ready; Progress added the veneer and tools that makes it fully in sync with open-source or agile BI as applied, say, to Big Data.

Thus, five years ago Progress Software would have been seen as an operational and decision-support database for an SMB, and a fast-moving local-level operational adjunct to enterprise BI in the Global 1000. Today, Progress Software has much less visibility in BI, and its connection to the latest BI technology is less visible; but if you look at the actual technology, Apama is indeed innovative and can deliver value-add across a wide range of enterprises. It only remains to ask, what’s the future of Progress Software as a supplier of event processing technology, and in general?

There are two parts to my answer. First, let me note the characteristic that Progress Software shares with just about every database company: it has a core loyal base of customers whose size may shrink, but whose tendency not to migrate away from the platform ensures that Progress Software will be extraordinarily long-lived. As in the past few years, database revenues may ebb over time; but few if any database companies in my 30 years of acquaintanceship with the industry see a massive collapse of their installed base. In the next 2-3 years, as sure as the sun rises, reasonable management plus this revenue flow will see Progress Software still standing (acquired or not) and still supporting Apama plus the infrastructure software like the Progress ESB that complements Apama.

However, there is little in the past three years, where revenues have been essentially flat, to suggest that Progress Software’s glory days will return, and that it will identify another new infrastructure technology to get strong growth started again. Moreover, Progress’ strength has never been that its solutions were a key part of the mainstream of Web innovation, and so there is no obvious reason to expect that Apama can, at a minimum, transition to form the core of cutting-edge open-source BI solutions. And that brings me to the second part of my answer.

I assert that these caveats almost certainly do not matter, in the next 2-3 years and probably further out. The reason is that Progress Software’s DNA may not be Web or open-source, but it is very definitely SMB-simple and flexible, with no vendor lock-in. Apama pre-acquisition would probably be a niche financial-market BI product. Apama plus Progress Software is a uniquely flexible and easy-to-use event processor that integrates with the rest of your architecture just fine, and it will stay that way. Progress Software’s new services prowess just ensures that the simplicity scales up to the largest of enterprises.

Potential Uses of DataRush for IT
The net of the contributions of both Apama technology and Progress Software’s “approach” to the Progress Apama solution is that, whether you are an SMB tackling BI for the first time or a large enterprise trying to become more agile in your querying of Big Data, Apama offers differentiated value-add, either stand-alone or as the complement to a large-vendor event processor and database architecture like IBM Streams and IBM Information Server. In the case of an SMB, the use case is straightforward: you can fit Apama with agile-BI efforts as a front end to catch urgent alerts implicit in Big Data, raw operational data, or ongoing analytics; and you can begin to develop EPM. Large enterprises typically have bottom-up Windows/desktop computing in parallel with the massive datacenter data warehouse: they can grow Apama with the evolution of that side of the enterprise architecture, while if appropriate also driving forward their top-down, datacenter-driven event processing projects using other vendors’ event-processing solutions, and easily integrating the two.

In either use case, the key to the most rapid possible success (I think) will be to identify on an ongoing basis the key targets for alerting and rapid decision-making, and to tie Apama as much as possible to historical data in order to allow the greatest “depth” of analytics at the point of data intercept. One good detection of a major customer about to dump you and immediate, effective reaction to prevent it will make the whole exercise more than worthwhile. And, remember, we are talking low-touch, very-low-TCO event processing here.

It is unusually hard to think of things that can go wrong with Progress Apama implementation. Services? Not as much needed, and Progress Software services that may be needed are already real-world-proven from more than 20 years of rave reviews from SMB and departmental clients, plus 5 years of driving similar technology into the division and LOB. Training? Again, the Progress Software track record is that even an untrained local-office manager can handle database maintenance – which is usually the biggest concern – and any developer can handle the drag-and-drop development tools. Integration? Progress Software has what I view as the standard set of adapters and gateways. Limits to scalability? Not if the Progress ESB is any guide. Let’s face it, IT implementation of Apama is very unlikely to be rocket science – and neither is gaining BI insights with it.

The Bottom Line for IT Buyers
At this point, an IT buyer should view Progress Software’s Apama as roughly equivalent to a “diamond in the rough.” It is the center of attention neither of BI, nor streaming technology, nor even sometimes of Progress Software itself. It suffers undeservedly from questions about Progress Software’s future, the future of event processing technology, whether Apama will continue to track BI technology, and a reputation as a high-end or specialized event processor. All it has going for it is that it is a more simple, more flexible, powerful event processing tool for a wide variety of use cases and scales, and should continue to deliver for the next 3-5 years, and almost certainly longer. And that should be plenty.

I have noted that Apama’s (and event processing’s) main value-add in BI is more in the risk area than in the top or bottom line (although some implementations, like EPM, do indeed impact revenues and costs). However, this is one technology that, when it succeeds, is really, really visible. Sell it to corporate as the latest technology fad if you like; the odds are that when the IT buyer acquires and then IT implements Apama, a big success story happens in the next year. And then you can concentrate on what’s really important: integration with the rest of your BI so that it all works optimally, in harmony.

The net-net for IT buyers, therefore, is to do a “reality check” on present-day event processing in BI, and then prepare a short list of event-processing software vendors to take the next step. I see no reason why, in most if not all cases, Progress Software’s Apama should not be on that Other BI “pre-short list.”