Tuesday, July 19, 2011

Will We All Speak IT?

At a recent teleconference, I heard the speaker first refer to “provisioning” a solution and then to people who would “on-board” that solution. It suddenly struck me that I was witnessing a new stage in the intrusion of language derived from computing into our daily lives.

Here’s how it used to go: we grew up with the rules of grammar and vocabulary as taught us in school, and as computer technology evolved, its new ideas and products used the words of, and fit neatly into, the English we were taught. A machine made a computation of a number, it computed, it was a computer. A piece of information in a computer, from the Latin, was a datum, plural data, stored in a data base, managed by a database management system.

In the same way, the jargon of computer techies, even when it spilled over into the population at large, was ultimately derived from ideas already in English. Bogosity – the quality of being bogus; bogon – a unit of bogosity. Misfeature – a combination of mistake and feature, a mistake that was touted by marketdroids (mindless marketers) as a feature.

ITSpeak 2.0
I first noticed things beginning to change in the late 1990s. In the 1980s, I had been frustrated as a purist by the universal tendency of my fellow programmers to refer to an example problem, an example case, an example screen, instead of a sample screen, as I had always been taught. Still, until the late 1990s, I never saw anyone else use “example” as an adjective; then marketing, and sometimes business blogs for a general audience, started to use “example” that way. However, even in the computing industry, there was strong purist resistance. I well remember the difficulties I had at Aberdeen Group persuading the editors that in computing, it was now “lifecycle”, not “life cycle”. Today, I can’t remember having seen “life cycle” in years.

In some ways, these tinkerings with basic English had a positive effect, I believe. Using “example” for “sample” is a good case in point: the meaning is clear from context, and it’s easier to use one word for the concept than learn two.

But the changes were not all for the good. I still remember some annoying marketer at Sybase, iirc, deciding in the late 1990s that from now on it was to be “database”, not “database management system”. The result was that users ever since are constantly confused as to whether they are talking about the software, or the data stored for use by that software – which I now have to always call the “data store” to make myself clear. In the same way, Enterprise Information Integration is now “data virtualization”, which captures only half the qualities of the software.

And, of course, with the advent of the Web IT words became far more ubiquitous, from blog to tweet. Sadly, these words have now become a measure of age, as each successive fad embeds its IT words into popular language, and we now divide generations into those that know what “to friend” means, and those who don’t.

ITSpeak Takes Over?
Even so, I didn’t see until now any clear indication that computer jargon was crowding out basic English words. But consider “provision”. Until very recently, a male was a “provider” who made enough money to put “provisions” on the table for the family. Now, IT has taken the word and abstracted it, to describe a general process of populating the empty shell of any new solution, and turned it from a noun into a verb. This major change in meaning is coming from IT, but it isn’t stopping there. Pretty soon, I expect to hear supermarkets start talking about “provisioning” their new stores, and then home builders and buyers start to talk about “provisioning” the new house with furniture.

The same goes for “on-board”. Like “friend” and “provision”, it’s a straightforward conversion of another part of speech to a verb. Like “provision”, it is a major switch in meaning that carries with it the notion of a process rather than an individual act. In the teleconference, it appeared to mean users carrying out the tasks of becoming part of a new IT solution themselves. But, again, I expect that soon employees will be expected to “on-board” themselves via “self-service portals”, and then students starting at college, and then what? Will we create new automated birthing centers where newborns will be expected to “on-board” themselves by responding to automated nipples? Will end-of-life hospices be referred to as “off-boarding centers?”

What If ITSpeak Does Take Over?
If we all starting talking ITSpeak – a language many of whose concepts originated in computing – is that good or bad? I believe that it’s way too early to tell. On the positive side, many of these words come from trying to distinguish more clearly between similar things, when the differences matter. The idea that a misfeature is not the same as a feature is important, and useful to us.

On the negative side, some historical richness of meaning may be lost. Always employing “utilize” instead of “use” (not really ITSpeak, but analogous) is not only unnecessarily lengthy, it also misses the importance in history of the distinction between “applying an object for a use for which it is designed” and “applying an object whether it helps in a task or not”. You utilize a Phillips screwdriver in following the directions for assembling a kid’s toy; you use a user’s manual for Microsoft Word even though it often doesn’t give you the answer you need.

No, my point here is that I think this represents a fundamental shift in our thinking, as we begin to see the world as IT folks do. At the least, this might mean that we think more of software-type abstractions and less of “legacy” physical objects, see life more in terms of processes and less in terms of interactions, and view others less in terms of irrationality and psychology and more in terms of categories and connections. So to maximize the chances of something good coming out of this, I think we ought to at least recognize that it is going on.

Will we all speak IT, all the time? Someday, quite possibly. Right now, it’s time to prepare to provision, so that we may on-board effectively.

Friday, July 8, 2011

The IBM Acquisition Game

Recently, I was contacted by a firm called Software Advice, which has a very interesting business model: pay-for-results advice on short lists for IT buying. They just posted a blog on “IBM M&A: Who’s Next”, and were interested in my thoughts. I took a look, and found it quite impressive; and therefore, in accordance with my philosophy of comforting the afflicted and afflicting the comfortable, I decided to pick nits about their conclusions. I believe that both their and my thoughts offer some potentially useful insights to IT buyers, not just about IBM, but about how quickly vendors are likely to deliver what users need in the next 1-2 years.

The reason it’s not only fun but instructive to play the IBM acquisition game is that it implicitly asks, given user needs over the next 1-2 years, what are the holes in IBM’s lineup to meet those needs that it should fill immediately? And that also allows us to ask, if they don’t fill those needs, will it come back to bite them, because someone else is likely to beat them to the punch? And then we can ask, will folks really want to use someone else besides IBM if IBM doesn’t supply this need – or is this something for which the IT buyer will have to “roll his or her own” at greater expense?

So let the game begin!

Historical Nits
Software Advice begins with a graphic nicely capturing the extent of IBM’s acquisitions over the last decade or so. The problem lies in the headings that split the acquisitions into “applications”, “infrastructure”, and “services”. You see, IBM has been firm in disclaiming any intention of getting into “applications”, and so most if not all of the acquisitions classified as “applications” are in fact what is usually called “infrastructure software.” That also means that almost identical infrastructure software is in one case classified as an application and in another as infrastructure – for example, the Rational software development toolset is counted as infrastructure software, but the Telelogic requirements management toolset, which is almost always used as the first step in the development process as part of a “lifecycle” software development toolset, is classified as an “application”. Hence it’s very easy to assume that IBM doesn’t need any applications acquisitions.

The interesting thing about this nit is that it raises the question: should IBM, at long last, go into the “apps business”, either on the business or consumer side? Yes, they’ve never needed to before, since until recently both Oracle and SAP (the dominant players in enterprise apps) have shown themselves willing to support all hardware vendors, but now that Oracle owns Sun and has shown it can play hardball with respect to HP Itanium, should IBM rethink that posture? Does the market now need a platform that it can be sure its enterprise or other business-critical applications will support?

The answer to that, I believe, depends on SAP. In other words, whatever the merits of other app vendors like Salesforce.com, the run-the-business applications of SAP are presently the main alternatives to Oracle Apps. If SAP remains a strong alternative, then IBM is entirely correct in continuing to keep its hands off enterprise application companies, reinforcing its image as less prone to vendor lock-in than Microsoft or Oracle.

And yet, I have to say, whether SAP will be a strong alternative remains an open question. SAP has made some major acquisitions of its own, like Business Objects and Sybase, which have taken it down the software stack with some quality infrastructure software. However, it is not yet clear that SAP can drive rapidly-changing database technology ahead fast enough to provide a long-run all-in-one enterprise-app or analytics alternative to Oracle Apps. The signs are very good: SAP appears to understand the importance of Sybase, and the potential of integrating its technologies with SAP’s present stack. Still, SAP has to execute that strategy.

I think it would make most sense for IBM to beef up its SAP application support with a smaller acquisition or two, this time of cross-database administrative tools that specialize in Sybase. Later, of course, if things get bad, IBM could always acquire SAP. In the meanwhile, the IT buyer should note that IBM and Oracle SAP support is a space to watch.

Strategic Investment Nits
Software Advice then goes on to identify general areas of future customer need where IBM may need to acquire companies. Their main focus – certainly a good one – is cloud administration. They also note – although with a much shorter analysis – IBM’s need to expand its analytics and BI offerings even further – and that makes sense too. Everyone, not just IBM, is scrambling to fill in the blanks and achieve fully automated hybrid-cloud deployment and administration.

However, I would disagree with their analysis of virtualization as a key area of acquisition. While VMWare continues to be an outstanding success, the pace of virtualization to a public cloud – the main lock-in for VMWare – remains quite slow. Private clouds in larger enterprises tend to be top-down, which means that IBM is doing quite well at driving its own virtualization software across the data center. I would argue that IBM has no need to acquire either VMWare or EMC either now or in the next two years – and a good reason to wait to see what happens as Oracle continues to compete with EMC more strongly in storage.

What might make sense, on the other hand, is for IBM to consider acquiring Red Hat. The two have been working together pretty effectively, and IBM needs to build up its open-source brand as a new market of tech-savvy open-source-oriented firms opens up. As long as it leaves the open-source culture of Red Hat in place, IBM can use Red Hat as an “early warning system” for changes in the new market – because that market cares less about VMWare vs. KVM and more about open-source-based services for cloud deployment.

My second nit regards mobile technology. It appears likely that the movement of mobile business workers towards having a laptop for some situations and a small-form-factor smartphone or tablet for others has reached flood stage, and needs to be addressed better. Sybase would have been a great entry point, but it’s not available now. Buying Apple would be fun to imagine, but seems impossible to achieve. Perhaps IBM might consider RIM. The value-add of Blackberry cell phones has always been in their business software, and while they are under threat in the consumer market, business users still find them appropriate. Here is an area of great user need where all vendors – not just IBM – fall short; so if IBM doesn’t do a good acquisition soon, IT buyers should anticipate a lot of “roll your own”.

My third nit concerns the whole area of BI/analytics. There seems to be a pervasive confusion of BI, analytics, and Big Data, as if they are the same thing. My short take on the differences is: BI is basic repeated reporting and querying plus ad-hoc or goal-oriented querying, both for corporate; analytics is ad-hoc or goal-oriented querying, not only for corporate but also embedded in other software across the organization (e.g., security and administrative analytics); Big Data is a wide range of new large-footprint data types, more usually on the Web, that provides insights into such new marketing topics as social media, and therefore typically complements BI with extra-organizational data. The result is that any good push to meet user needs is going to need to tackle all three areas.

As I noted in a previous blog post, what’s users need in all three areas is some combination of scaling and user friendliness, especially for the burgeoning SMB BI market. It’s hard to buy or create user friendliness – the BI market still has a ways to go in this area. However, there are ways that IBM could improve its scalability. For one thing, Netezza and the new IBM z appliance have columnar database technology that’s too tied to a particular appliance. It’s not clear just how fast IBM will move into this area, but Amazon’s investment in ParAccel reminds us that there are still interesting columnar database suppliers out there.

On the Big Data side, users must also consider integrating BI with file-system-stored Web data such as that accessed via Hadoop. There are quite a few NoSQL open-source efforts that may be worth productizing and integrating with DB2 or a columnar database. Again, this is an area where all vendors – not just IBM – need to do more to make the path to combined BI/analytics/Big Data clear. In the meanwhile, IT buyers should think carefully about buying from only one database vendor, because until one of them shows they have the full Big Data story there is no guarantee that any of them will not fall short of what users need – and past experience suggests that database lock-in is about as locked in as you can get.

Endgame
Having played Software Advice’s IBM Acquisition Game, I draw three conclusions from it. First, IBM is in a surprisingly strong position going forward. There is no obvious hole in its solution lineup that immediately threatens the company, and that it cannot fix by careful re-tuning of the same strategies it has had up to now. And that’s good news for IT buyers.

Which leads me to conclusion two: there’s still enough choice in the market. We have seen a lot of acquisitions, not just from IBM but from other major vendors, in the last decade; but the fact that there are still smaller companies out there to plug holes for IBM and others means that IT buyers can still find a way to stitch together a solution where one’s favored vendor doesn’t quite cover all needs.

And that leads to conclusion three: despite the hype, users are still a long way from taking full advantage of mobile, cloud, or analytics/Big Data. This may well be a transition as slow and incomplete as the one in the early 2000s to service-oriented architectures – and don’t get me started about Business Process Integration. In fact, it might be a good idea to play the Acquisition Game with other vendors on your short lists – and then see what those vendors do in the real world to cover the holes you find, before committing irrevocably and totally to one of them. That’s not to say you shouldn’t press ahead with all deliberate speed, as your competitors will be doing -- but cover your bets.

Wednesday, June 29, 2011

Big Data, MapReduce, Hadoop, NoSQL: Pay No Attention to the Relational Technology Behind the Curtain

One of the more interesting features of vendors’ recent marketing push to sell BI and analytics is the emphasis on the notion of Big Data, often associated with NoSQL, Google MapReduce, and Apache Hadoop – without a clear explanation of what these are, and where they are useful. It is as if we were back in the days of “checklist marketing”, where the aim of a vendor like IBM or Oracle was to convince you that if competitors’ products didn’t support a long list of features, that those competitors would not provide you with the cradle-to-grave support you needed to survive computing’s fast-moving technology. As it turned out, many of those features were unnecessary in the short run, and a waste of money in the long run; remember rules-based AI? Or so-called standard UNIX? The technology in those features was later to be used quite effectively in other, more valuable pieces of software, but the value-add of the feature itself turned out to be illusory.

As it turns out, we are not back in those days, and Big Data via Hadoop and NoSQL does indeed have a part to play in scaling Web data. However, I find that IT buyer misunderstandings of these concepts may indeed lead to much wasted money, not to mention serious downtime. These misunderstandings stem from a common source: marketing’s failure to explain how Big Data relates to the relational databases that have fueled almost all data analysis and data-management scaling for the last 25 years. It resembles the scene in Wizard of Oz where a small man, trying to sell himself as a powerful wizard by manipulating stage machines from behind a curtain, becomes so wrapped up in the production that when someone notes “There’s a man behind the curtain” the man shouts “Pay no attention to the man behind the curtain!” In this case, marketers are shouting about the virtues of Big Data related to new data management tools and “NoSQL” that they fail to note the extent to which relational technology is complementary to, necessary to, or simply the basis of, the new features.

So here is my understanding of the present state of the art in Big Data, and the ways in which IT buyers should and should not seek to use it as an extension of their present (relational) BI and information management capabilities. As it turns out, when we understand both the relational technology behind the curtain and the ways it has been extended, we can do a much better job of applying Big Data to long-term IT tasks.

NoSQL or NoREL?
The best way to understand the place of Hadoop in the computing universe is to view the history of data processing as a constant battle between parallelism and concurrency. Think of the database as a data store plus a protective layer of software that is constantly being bombarded by transactions – and often, another transaction on a piece of data arrives before the first is finished. To handle all the transactions, databases have two choices at each stage in computation: parallelism, in which two transactions are literally being processed at the same time, and concurrency, in which a processor switches between the two rapidly in the middle of the transaction. Pure parallelism is obviously faster; but to avoid inconsistencies in the results of the transaction, you often need coordinating software, and that coordinating software is hard to operate in parallel, because it involves frequent communication between the parallel “threads” of the two transactions.

At a global level (like that of the Internet) the choice now translates into a choice between “distributed” and “scale-up” single-system processing. As it happens, back in graduate school I did a calculation of the relative performance merits of tree networks of microcomputers versus machines with a fixed number of parallel processors, which provides some general rules. There are two key factors that are relevant here: “data locality” and “number of connections used” – which means that you can get away with parallelism if, say, you can operate on a small chunk of the overall data store on each node, and if you don’t have to coordinate too many nodes at one time.

Enter the problems of cost and scalability. The server farms that grew like Topsy during Web 1.0 had hundreds and thousands of PC-like servers that were set up to handle transactions in parallel. This had obvious cost advantages, since PCs were far cheaper; but data locality was a problem in trying to scale, since even when data was partitioned correctly in the beginning between clusters of PCs, over time data copies and data links proliferated, requiring more and more coordination. Meanwhile, in the High Performance Computing (HPC) area, grids of PC-type small machines operating in parallel found that scaling required all sorts of caching and coordination “tricks”, even when, by choosing the transaction type carefully, the user could minimize the need for coordination.

For certain problems, however, relational databases designed for “scale-up” systems and structured data did even less well. For indexing and serving massive amounts of “rich-text” (text plus graphics, audio, and video) data like Facebook pages, for streaming media, and of course for HPC, a relational database would insist on careful consistency between data copies in a distributed configuration, and so could not squeeze the last ounce of parallelism out of these transaction streams. And so, to squeeze costs to a minimum, and to maximize the parallelism of these types of transactions, Google, the open source movement, and various others turned to MapReduce, Hadoop, and various other non-relational approaches.

These efforts combined open-source software, typically related to Apache, large amounts of small or PC-type servers, and a loosening of consistency constraints on the distributed transactions – an approach called eventual consistency. The basic idea was to minimize coordination by identifying types of transactions where it didn’t matter if some users got “old” rather than the latest data, or it didn’t matter if some users got an answer but others didn’t. As a communication from Pervasive Software about an upcoming conference shows, a study of one implementation finds 60 instances of unexpected unavailability “interruptions” in 500 days – certainly not up to the standards of the typical business-critical operational database, but also not an overriding concern to users.

The eventual consistency part of this overall effort has sometimes been called NoSQL. However, Wikipedia notes that in fact it might correctly be called NoREL, meaning “for situations where relational is not appropriate.” In other words, Hadoop and the like by no means exclude all relational technology, and many of them concede that relational “scale-up” databases are more appropriate in some cases even within the broad overall category of Big Data (i.e., rich-text Web data and HPC data). And, indeed, some implementations provide extended-SQL or SQL-like interfaces to these non-relational databases.

Where Are the Boundaries?
The most popular “spearhead” of Big Data, right now, appears to be Hadoop. As noted, it provides a distributed file system “veneer” to MapReduce for data-intensive applications (including Hadoop Common that divides nodes into a master coordinator and slave task executors for file-data access, and Hadoop Distributed File System [HDFS] for clustering multiple machines), and therefore allows parallel scaling of transactions against rich-text data such as some social-media data. It operates by dividing a “task” into “sub-tasks” that it hands out redundantly to back-end servers, which all operate in parallel (conceptually, at least) on a common data store.

As it turns out, there are also limits even on Hadoop’s eventual-consistency type of parallelism. In particular, it now appears that the metadata that supports recombination of the results of “sub-tasks” must itself be “federated” across multiple nodes, for both availability and scalability purposes. And Pervasive Software notes that its own investigations show that using multiple-core “scale-up” nodes for the sub-tasks improves performance compared to proliferating yet more distributed single-processor PC servers. In other words, the most scalable system, even in Big Data territory, is one that combines strict and eventual consistency, parallelism and concurrency, distributed and scale-up single-system architectures, and NoSQL and relational technology.

Solutions like Hadoop are effectively out there “in the cloud” and therefore outside the enterprise’s data centers. Thus, there are fixed and probably permanent physical and organizational boundaries between IT’s data stores and those serviced by Hadoop. Moreover, it should be apparent from the above that existing BI and analytics systems will not suddenly convert to Hadoop files and access mechanisms, nor will “mini-Hadoops” suddenly spring up inside the corporate firewall and create havoc with enterprise data governance. The use cases are too different.

The remaining boundaries – the ones that should matter to IT buyers – are those between existing relational BI and analytics databases and data stores and Hadoop’s file system and files. And here is where “eventual consistency” really matters. The enterprise cannot treat this data as just another BI data source. It differs fundamentally in that the enterprise can be far less sure that the data is up to date – or even available at all times. So scheduled reporting or business-critical computing based on this data is much more difficult to pull off.

On the other hand, this is data that would otherwise be unavailable – and because of the low-cost approach to building the solution, should be exceptionally low-cost to access. However, pointing the raw data at existing BI tools is like pointing a fire hose at your mouth. The savvy IT organization needs to have plans in place to filter the data before it begins to access it.

The Long-Run Bottom LineThe impression given by marketers is that Hadoop and its ilk are required for Big Data, where Big Data is more broadly defined as most Web-based semi-structured and unstructured data. If that is your impression, I believe it to be untrue. Instead, handling Big Data is likely to require a careful mix of relational and non-relational, data-center and extra-enterprise BI, with relational in-enterprise BI taking the lead role. And as the limits to parallel scalability of Hadoop and the like become more evident, the use of SQL-like interfaces and relational databases within Big Data use cases will become more frequent, not less.

Therefore, I believe that Hadoop and its brand of Big Data will always remain a useful but not business-critical adjunct to an overall BI and information management strategy. Instead, users should anticipate that it will take its place alongside relational access to other types of Big Data, and that the key to IT success in Big Data BI will be in intermixing the two in the proper proportions, and with the proper security mechanisms. Hadoop, MapReduce, NoSQL, and Big Data, they’re all useful – but only if you pay attention to the relational technology behind the curtain.

Pentaho and Open Source BI: The New SMB

On Monday, Pentaho, an open source BI vendor, announced Pentaho BI 4.0, its new release of its “agile BI” tool. To understand the power and usefulness of Pentaho, you should understand the fundamental ways in which the markets that we loosely call SMB have changed over the last 10 years.

First, a review. Until the early 1990s, it was a truism that computer companies in the long run would need to sell to central IT at large enterprises, eventually – else the urge of CIOs to standardize on one software and hardware vendor would favor larger players with existing toeholds in central IT. This was particularly true in databases, where Oracle sought to recreate the “nobody ever got fired for buying IBM” hardware mentality of the 1970s in software stacks. It was not until the mid-1990s that companies such as Progress Software and Sybase (with its iAnywhere line) showed that databases delivering near-lights-out administration could survive the Oracle onslaught. Moreover, companies like Microsoft showed that software aimed at the SMB could over time accumulate and force its way into central IT – not only Windows, Word, and Excel, but also SQL Server.

As companies such as IBM discovered with the bursting of the Internet bubble, this “SMB” market was surprisingly large. Even better, it was counter-cyclical: when large enterprises whose IT was a major part of corporate spend cut IT budgets dramatically, SMBs kept right on paying the yearly license fees for the apps on which they ran, which in turn hid the brand on the database or app server. Above all, it was not driven by brand or standards-based spending, nor even solely by economies of scale in cost.

In fact, the SMB buyer was and is distinctly and permanently different from the large-enterprise IT buyer. Concern for costs may be heightened, yes; but also the need for simplified user interfaces and administration that a non-techie can handle. A database like Pervasive could be run by the executive at a car dealership, who would simply press a button to run backup on his or her way out on the weekend, or not even that. The ability to fine-tune for maximum performance is far less important than the avoidance of constant parameter tuning. The ability to cut hardware costs by placing apps in a central location matters much less than having desktop storage to work on when the server goes down.

But in the early 2000s, just as larger vendors were beginning to wake up to the potential of this SMB market, a new breed of SMB emerged. This Web-focused SMB was and is tech-savvy, because using the Web more effectively is how it makes its money. Therefore, the old approach of Microsoft and Sybase when they were wannabes – provide crude APIs and let the customer do the rest – was exactly what this SMB wanted. And, again, this SMB was not just the smaller-sized firm, but also the skunk works and innovation center of the larger enterprise.

It is this new type of SMB that is the sweet spot of open source software in general, and open source BI in particular. Open source has created a massive “movement” of external programmers that have moved steadily up the software stack from Linux to BI, and in the process created new kludges that turn out to be surprisingly scalable: MapReduce, Hadoop, noSQL, and Pentaho being only the latest examples. The new SMB is a heavy user of open source software in general, because the new open source software costs nothing, fits the skills and Web needs of the SMB, and allows immediate implementation of crude solutions plus scalability supplied by the evolution of the software itself. Within a very few years, many users, rightly or wrongly, were swearing that MySQL was outscaling Oracle.

Translating Pentaho BI 4.0
The new features in Pentaho BI can be simply put, because the details simply show that they deliver what they promise:

· Simple, powerful interactive reporting – which apparently tends to be used more for ad-hoc reporting that the traditional enterprise reporting, but can do either;
· A more “usable” and customizable user interface with the usual Web “sizzle”;
· Data discovery “exploration” enhancements such as new charts for better data visualization.

These sit atop a BI tool that distinguishes itself by “data integration” that handles an exceptional number of input data warehouses and data stores for inhaling to a temporary “data mart” for each use case.

With these features, Pentaho BI, I believe, is valuable especially to the new type of SMB. For the content-free buzz word “agile BI”, read “it lets your techies attach quickly to your existing databases as well as Big Data out there on the Web, and then makes it easy for you to figure out how to dig deeper as a technically-minded user who is not a data-mining expert.” Above all, Pentaho has the usual open source model, so it’s making its money by services and support – allowing the new SMB to decide exactly how much to spend. Note also Pentaho’s alliance not merely with the usual cloud open source suspects like Red Hat but also with database vendors with strong BI-performance technology such as Vertica.

The BI Bottom Line
No BI vendor is guaranteed a leadership position in cloud BI these days – the field is moving that fast. However, Pentaho is clearly well suited to the new SMB, and also understands the importance of user interfaces, simplicity for the administrator, ad hoc querying and reporting, and rapid implementation to both new and old SMBs.

Pentaho therefore deserves a closer look by new-SMB IT buyers, either as a cloud supplement to existing BI or as the core of low-cost, fast-growing Web-focused BI. And, remember, these have their counterparts in large enterprises – so those should take a look as well. Sooner than I expected, open source BI is proving its worth.

Wednesday, June 22, 2011

This Made Me Sick To My Stomach

I just finished reading the report summary of the International Earth system expert workshop on ocean stresses and impacts, released on Monday. Here is what I think the headline of that report should have been:

The Oceans Are Mostly DEAD Unless We Reduce Carbon Emissions Drastically AND Set Up a Global Police Over ALL Ocean Uses NOW

I won’t bother with the evidence of cascading destruction, with more to come, and the explanation of how much of what we do now that affects the oceans reinforces that cascade. That has been covered somewhat in Joe Romm’s blog, www.climateprogress.com. I will simply note that the end point will be a mass ocean species destruction comparable to any in the past, plus massive ocean acid “dead zones” where nothing can live and a time to recover in the thousands of years. One species that might survive is jellyfish – and it has “low nutritional value,” i.e., you can’t live on jellyfish.

The one thing that no one seems to be covering is what they say we should do to avoid this. They say that everyone on Earth must stop all ocean misuses now, and to do this the UN should set up a global enforcement body. The burden of proof will be on all ocean users – yes, that includes ocean liners and drillers for oil, gas, and minerals – to show that their next use will not be harmful, else they can’t do it. Contributions to the body would be mandatory, and it would have jurisdiction over the “High Seas” that aren’t the property of particular nations, but obviously it will affect waters that are now said to be the property of particular nations, as well as fisheries.

If you want to go fish, get permission from the global enforcement body. If you want to ship components from abroad, get permission from the global enforcement body. If you want to drill in the Arctic now that it’s getting warmer and less icy, get permission. If you dump fertilizer and waste into rivers and it’s washed out to sea, the commission will be after you. And the commission’s key criteria will be: Does this add to the carbon footprint? Does this make a dead ocean more likely? Is this a sustainable use?

As far as I can see, the only reason the workshop would recommend such a thing is that the situation is that serious. And it’s serious not only because we lose seafood, but because the ocean will reach its capacity for absorbing the excess carbon we’re dumping in the atmosphere, and then global warming on land will get worse, faster than we expect even now – leading to faster sea rise and more massive storms that “salt” estuaries that product a significant proportion of the world’s food, more droughts over much of the world that desertify another major proportion of the world’s food, and possibly to “toxic blooms” in the waters next to the land that periodically release toxic gases that kill those living on the shore.

Just thinking about the ocean, near which I have lived for most of my life, being dead makes me sick to my stomach.

Sunday, June 12, 2011

In the End, Godel Has Won

This post was originally written last fall, and set aside as being too speculative. I felt that there was too little evidence to back up my idea that “accepting limits” would pay off in business.

Since then, however, the Spring 2011 edition of MIT Sloan Management Review has landed on my desk. In it, a new “sustainability” study shows that “embracers” are delivering exceptional comparative advantage, and that a key characteristic of “embracers” is that they view “sustainability” as a culture to be “wired into the business” – “it’s the mindset”, says Bowman of Duke Energy. According to Wikipedia, the term “sustainability” itself is fundamentally about accepting limits, including environmental “carrying capacity” limits, energy limits, and limits in which use rates don’t exceed regeneration rates.

This attitude is in stark contrast to the attitude pervading much of human history. I myself have grown up in a world in which one of the fundamental assumptions, one of the fundamental guides to behavior, is that it is possible to do anything. The motto of the Seabees in World War II, I believe, was “The difficult we do immediately; the impossible takes a little longer.” Over and over, we have believed adjustments in the market, inventions and advances, daring to try something else, an all-out effort, something, anything, can fix any problem.

In mathematics, they, too, believed at the turn of the century that any problem was solvable: that any truth of any consistent, infinite mathematical system could be proved. And then Kurt Godel came along and showed that in every such system, either you could not prove all truths or you could also prove false things, one or the other. And over the next thirty years, mathematics applied to computing showed that some problems were unsolvable, and others had a fundamental lower limit on the time taken to solve the problem that meant that they could not be solved before the universe ended. By accepting these limits, mathematics and programming have flourished.

This mindset is fundamentally different from the “anything is possible” mindset. It says to work smarter, not harder, by not wasting your time on the unachievable. It says to identify the highly improbable up front and spend most of your time on solutions that don’t involve that improbability. It says, as agile programming does, that we should focus on changing our solutions as we find out these improbabilities and impossibilities, rather than piling on patch after patch. It also says, as agile programming does, that while by any short-run calculation the results of this mindset might seem worse than the results of the “anything is possible” mindset, over the long run – and frequently over the medium term – it will produce better results.

It seems more and more apparent to me that we have finally reached the point where the “anything is possible” approach is costing us dearly. I am speaking specifically about climate change – one key driver for the sustainability movement. The more I become familiar with the overwhelming scientific evidence for massive human-caused climate change and the increasing inevitability of at least some major costs of that change in every locality and country of the globe, the more I realize that an “anything is possible” mentality is a fundamental cause of most people’s failure to respond adequately so far, and a clear predictor of future failure.

Let me be more specific: as noted in the UN scientific conferences and recent additional data, “business as usual” is leading us to a carbon dioxide concentration of 1000 ppm in the atmosphere, of which about 450 ppm or 150-200 ppm over the natural amount is already “baked in”. This will result, at minimum, in global increases in temperature of 5-10 degrees Fahrenheit, which will result, among other things, in order-of-magnitude increases in the damage caused by extreme weather events, the extinction of many ecosystems supporting existing urban and rural populations – because many of these ecosystems are blocked from moving north or south by paved human habitations – so that food and shelter production must both change their location and find new ways to deliver to new locations, movement of all populations from locations on seacoasts up to 20 feet above existing sea level, and adjustment of a large proportion of heating and cooling systems to a new mix of the two – not to mention drought, famine, and economic stress. And these are just the effects over the next 60 or so years.

Adjusting to this will place additional costs on everyone, very possibly similar to a 10% tax yearly on every individual and business in every country for the next 50 years, no matter how wealthy or adept. Continuing “business as usual” for another 30 years would result in a similar, almost equally costly additional adjustment.

Our response to this so far has been in the finest tradition of “anything is possible”. We search for technological fixes under the belief that they will solve the problem, since they appear to have done so before. Most of us – except the embracers – assume that existing business incentives, focused on cutting costs – but these costs have not yet occurred – will somehow respond years before the impact begins to be felt. (Embracers, by the way, actively seek out new metrics to capture things like carbon emissions’ negative effects) We are skeptical and suspicious, since those who have predicted doom before, for whatever reason, have generally seemed to have turned out to be wrong. We hide our heads in the sand, because we have too much else to do and concerns that seem more immediate. We are distracted by possible fixes, and by their flaws.

The “embrace limits” mindset for climate changes makes one simple change: accept steady absolute reductions in carbon emissions as a limit. For example, every business, every country, every region, every county accepts that every year, its emissions are to be reduced by 1% in that year. If a business, that business also accepts that its products’ emissions are to be reduced by 1% in that year, no matter how successful the year has been. If a locality does better one year, it still is expected not to increase emissions the next year. If a country rejects this idea, investments from conforming countries are reduced by 1% each year, and products accepted from that country are expected to comply.

But this is a crude, blunt-force suggested application of “embrace limits”. There are all sorts of other applications. Investors will no longer invest in equities that seem to promise 2% long-term returns above historical norms, and will limit the amount of their capital invested in “bets,” because those investments are overwhelmingly likely to be con jobs. Project managers will no longer use metrics like time to deployment, but rather “time to value” and “agility”, because there is a strong possibility that during the project, the team will discover a limit and need to change its objective.

Because, fundamentally, climate change is a final, clear signal that Godel has won. Whether we accept limits or not, they are there; and the less we accept them and use them to work smarter, the more it costs us.

Tuesday, May 17, 2011

HP, IBM, and So On: It's Not the Personalities, People!

I note from various sources that some VSPs (Very Serious People, to borrow an acronym from Paul Krugman) are now raising questions about HP’s financials in the wake of Mark Hurd’s departure for Oracle. To cherrypick some quotes: “They need to .. regain investor confidence”; “HP is in a difficult situation”; “It sounds like … Hurd took too many costs out of [the services] business”; “HP … now … are known for inconsistency … It could become a value trap.” And, of course, there are comparisons with IBM, Dell, software vendors like Oracle, and so on.

I am certainly not an unalloyed HP booster. In fact, I have made many unflattering comparisons of HP with IBM myself over the years. However, I disagree with the apocalyptic tone of these pronouncements. In fact, I will stick out my neck and predict that HP will not implode over the next 3 years, and it will not fall behind IBM in revenues either, barring a truly epochal acquisition by IBM. I believe that these VSPs are placing too much emphasis on upcoming strategies bearing the imprint of personalities like HP’s Leo Aptheker and IBM’s Sam Palmisano, and not enough emphasis on the existing positioning of IBM, Dell, HP, Microsoft, Oracle, and Apple.

Let’s Start With the Negative!

So what are these problems that I have criticized HP for? Well, let’s start with its solution portfolio. Of the major computer vendors, HP may be the closest to a conglomerate – and that’s not a good thing. Let’s see, it has a printer/all-in-one company, a PC company, one or two server companies (including Tandem), a business/IT services/outsourcing company, and even, if you want to stretch a point, an administrative software utility company (the old SystemView) with some more recent software (Mercury Interactive testing) attached. Moreover, because HP has until very recently not tried very hard to stitch these together either as solutions or as a software/hardware stack, they are not as integrated as others – strikingly, not as integrated as IBM, which was once known for announcing global solutions whose components turned out to be in the early stages of learning to talk to each other. At first glance, HP’s endowments seem impressive; closer up, these seem, as someone once said in another context, like cats fighting in a burlap bag.

Moreover, HP, unlike any of the other companies I have mentioned except Dell, simply does not have software in its DNA. Back in the early 1990s, a Harvard Business Review article asserted that hardware companies at the time would suffer unless they became primarily services companies; I asserted then, and I assert now, that they also should become software companies.

I believe that this lack of software solutions and development personnel has several bad effects that have decreased HP’s revenues and profits by at least 20% over the last 20 years. Software development connects you with the open source community, the consumer market, and the latest technologies that impinge on computer vendors quite effectively. It allows your services arm to offer more leading-edge services, rather than trying to customize others’ software in a quick and dirty fashion for one particular services engagement. And, in the end, it moves hardware development ahead faster, as it focuses chip development on major real-world workloads that your software supports. Moreover, as IBM itself has proven, even if investment in software doesn’t pay off immediately, eventually you get it right.

A third, more recent problem, does relate to Mark Hurd’s cost focus – although the same might be said for IBM. A truism of business strategy proved by the problems created by CFO dominance at US car companies in the 1980s and 1990s is that too long a focus on the financials rather than product innovation costs a company dearly. It is quite possible that HP has eaten its innovation seed corn in the process of turning into a “consistent” money maker.

Finally, HP has in the past had a tendency in its hardware products to be “the nice alternative”: not locking you in or like Sun or Microsoft, willing to provide a platform for Oracle and Microsoft databases, open to anyone’s middleware. Whatever the merits of that approach, it creates a perception among customers that HP is not leading-edge in the sense that Apple, or even Microsoft and Oracle, are. Twenty-one years ago, in my first HP briefing, famous analyst Nina Lytton showed up in a brilliant pink outfit and immediately announced that HP’s strategy reminded her of a “great pink cloud.” That sort of rosy but not clear-cut presentation of one’s strategy and future plans does not create the sort of excitement among customers that Steve Jobs’ iPhone and iPad announcements, or even IBM’s display of Watson, do.

It Doesn’t Matter

And yet, when we look at HP vs. IBM in the longer run – from 1990, when I started as an analyst, to now – the ongoing success of HP is striking. At the start, IBM’s yearly revenues were in the $80s billion, and HP’s perhaps a quarter as much. Today, HP’s revenues are perhaps 1/3 greater, IBM at around a $100B run rate and HP perhaps at $135B. Some of that HP growth can be attributed to acquisitions; but a lot of it comes from growth of its core business and its acquisitions. To put it another way, IBM has been very successful at growing its profit margin; HP has been very successful at growing.

And growth at that scale is not easy. Companies have been trying to knock IBM off its Number 1 perch in revenues since the 1960s, and only HP has succeeded. Nobody else is in striking distance yet – Microsoft is at a $70B run rate, apparently, with seemingly no prospects of exceeding $100B in the next couple of years.

The reason, I think, is HP’s acquisition of Compaq back in the 1990s. Since then, having beaten back Dell’s challenge, HP is in a very strong position in the PC-form-factor scale-out markets. Despite recent apparent gains by System x, IBM focuses on the business market, and all of the other vendors mentioned above do not compete in PC hardware. Moreover, the PC market aligns HP with Intel and Microsoft, and thereby is relatively well protected from IBM’s POWER chips or even Oracle/Sun’s SPARC chipset, whatever life that has left in it (there is still no sign that AMD threatens Intel’s chip dominance significantly).

So let HP’s scale-up servers and storage falter in technology (e.g., the Itanium bet) relative to IBM and EMC, if they do; with the steady decrease in market share by Sun, HP is, and will in the short term remain, the IBM alternative in this market. Let Dell and IBM’s System x tout their business Linux scale-out prowess; the prevalence of existing scale-out PCs in public clouds and Microsoft LOBs means that HP is well positioned to handle competition in those areas over the next couple of years.
And who else but IBM can attack HP? Oracle may talk big, but Sun’s market share appears to be shrinking, and 15 years of Larry Ellison talking about the virtual desktop and Oracle Database handling your filesystem have failed to make a dent in Windows, much less Wintel. Microsoft has no need to move into hardware, and apparently no desire. Apple appears to be playing a different sport altogether.

In fact, the only serious threat to HP over the short term is any major movement of consumers off PCs and laptops as they move to smartphones and tablets. Here again, I think, analysts are too apocalyptic. Yes, iPhones can handle an astonishing range of consumer tasks, but not as easily or in as sophisticated a fashion as PCs, and users still continue to want to create and organize personal stores of photos etc. as well as share them – something the smartphone does not yet do. Meanwhile, the tablet offers the small form factor and attractive user interface that today’s laptop does not; but it is more likely that the tablet will acquire PC features, than that it will morph into an iPhone.

And Whither IBM?

In fact, an interesting question, given IBM’s status as the most direct competitor of HP, is whether IBM can begin to speed up its revenue growth. IBM has been delivering strong financials for almost 20 years, while talking a good game about innovation. In fact, I would say that they have indeed been innovative in some areas – but not enough yet to grow their revenues fast. Will the big innovation be green technology? The cloud? Analytics? Because, let’s face it, the only two things that recently have delivered big revenue gains are cell phones and Web 2.0/social media – and Apple, Google, and Facebook are the ones reaping the most revenues from these, not IBM.

In fact, as I have argued, IBM can do quite well with its present strong position in scale-up, but it cannot dominate the business side of computer markets when HP, Microsoft, and Intel have such a strong position in scale-out, nor can it match HP in consumer markets – and these affect business sales.

User Bottom Line: Don’t Panic, Do Buy Both

It would be nice, wouldn’t it, to be back in the old days when no one ever got fired for buying IBM systems, or Oracle databases? Well, those days are gone forever, and no blunder or inspired move, by Aptheker, Palmisano, Hurd, Ellison, Ballmer, Dell, or Jobs, will bring them back.

Given that, the smart IT buyer will acquire a little of each, in the areas in which each is best. It is true, for example, that IBM has exceptional services scope that allows effective integration – including integration of scale-out technology from Microsoft, Intel, and HP, or for that matter (System x) from IBM itself. This “mixed” enterprise architecture is the New Normal; vendor lock-in or a tide of Web innovation fueled by an Oracle and a Sun is so 1990s.

It is said that when Mary Queen of Scots wed the King of France, she was saluted with: “Let others wage war; let you, happy Scotland, bear children” (it’s better in Latin). Let the VSPs and apocalyptic analysts assert that vendor personalities waging war should affect your buying decision; you, happy CIO, should buy products from any of the vendors mentioned above, without worrying that a vendor is about to go belly-up in two seconds. And the vendors that have the greatest ability to integrate, like IBM and HP, will do quite well if you do.