Recently, I read a Paul Krugman NY Times column on income inequality, which referenced a NY Times article on income inequality, which referenced a Pew Research Center study on income mobility over generations. The NY Times article stated flatly that the Pew study found that 81% of Americans have more income than their parents. I read the first part of the study carefully, and it did indeed state that most sons of parents in the study, whether white or African-American, bottom or middle or top third of parental income, earned more than their parents did. And then I read even more carefully, and realized that the data absolutely did not support a conclusion that today’s American sons earn more than their parents did.
What went wrong? Well the study took a longitudinal study of families whose sons were between 0-18 years of age in 1967-1971, and compared the family income of the parents in 1967-1971 to the family income of the sons in 1995-2002 (omitting a couple of years). There were three basic problems with the analysis. First, government data shows that for the bottom third of family incomes, family income grew from 1967 to 1979, decreased slightly until 1994, grew again until 2001, and decreased to below the 1979 level by 2010. For the middle third, there was a similar trend, except that family incomes are now about the same or slightly below the 1979 level. For the upper third (and especially for the upper 1%), family income has grown consistently and, by 2010, substantially over 1967-1971. So the choice of the two “snapshot” time periods maximized the growth of income between parents and sons in all three income strata. Thus, it is likely that had the period been, say, 1980-1984 vs. 2006-2010, far fewer sons would have increased their family income compared to their parents.
The second problem with the study was the interval of the two “snapshots”. We know from demographic data that people were having kids, on average, earlier, in the 1950s and 1960s, and so it is reasonable to suppose that those sons who were 0-18 in 1967-1971 typically had parents that were 21–49, and most frequently about 35, while the sons in turn during 1995-2002 would be 24-55, and most frequently around 40. The reason this matters is that government data shows that families’ earnings trend steadily upwards from 20 onwards, and reach their peak from 45-55. In other words, the time period chosen exaggerated the income earned by the sons by pushing more of them into a peak earnings period.
The third problem with the study’s statistical approach is that it took “family income” as equivalent to “personal income.” Back in 1967-1971, across all three strata, less than one-third of women worked. According to the latest Census data, perhaps 80% as many women as men work, and that was pretty much true in the 1995-2002 period. These, in turn, have been earning perhaps 80% as much as men. So, especially in the bottom and middle thirds, women contributed less than 10% of average “family income” in the 1967-1971 period, and about 40% of “family income” in the 1995-2002 period. If we are really comparing apples to apples, we have to say that if we compare fathers to sons, it is clear that any upward trend in income is far less frequent. Now, the study notes that the “family income” is converted to personal income by being “family-size adjusted in all analyses”; but all this does is exaggerate things even further, because family sizes were slightly smaller in 1995-2002 (and now) than in 1967-1971.
One caveat: I was unable to access the Appendix to the study, which explained Pew’s methodology in greater detail. It is always possible that they dealt with these problems to some extent by further statistical tweaks. However, I view that as pretty unlikely, since these considerations are so important that they should have been noted in some way in the main paper.
Now, it is important to keep in mind that I am not a statistics “rocket scientist.” All it took me to figure this one out was a little ongoing digging on the topic of income inequality, and a careful lay-person reading of the methodology section of the Pew paper. The problem here is not that Pew was “lying with statistics”, because the facts were right there in the front of their study report. The real problem is that the so-called journalist of the NY Times apparently didn’t even bother to read that section carefully, much less do a little additional research which would have called the Times’ “81% of Americans” into further question.
So, as that fellow in the insurance commercials would say, what have we learned here? Well, first of all, statistical nits matter. I suspect that when the dust settles, we will find that less than half of all American males are presently earning as much, in real terms, as their fathers did (and the women aren’t a slam dunk either, since much of the surge in their employment and wages happened by the late 1980s). Even if that isn’t true, there’s no way the figure is anywhere near 81%. You need to consider statistical nits like the ones I have cited to convert a statistical study into a realistic picture of what’s going on in the real world.
Second, and equally important, you can’t trust any old reporter to do it for you, no matter how prestigious the name of their institution. You at least have to make a stab at the statistical nits yourself – or you’ll wind up believing what just isn’t so. Thank heavens for the Web, so that we can begin to check those statistical nits. Thank heavens for the Web, which gives us pointers to data that put those statistical nits in context. If we fail to do so, then by all means blame the knucklehead at the NY Times – but also blame ourselves.
Monday, January 9, 2012
Friday, January 6, 2012
The Other BI: Composite CIS/Studio and Agile BI
This blog post highlights a software company and technology that I view as potentially useful to organizations investing in business intelligence (BI) and analytics in the next few years. Note that, in my opinion, this company and solution are not typically “top of the mind” when we talk about BI today.
The Importance of the Composite Software Technology to BI
Today’s enthusiasm for “agile BI” that often consists of slapdash applications of SCRUM to rapid roll-out of new analytics applets obscures the fact that business agility, not to mention New Product Development (NPD) agility, involves applying more of agile methodologies than a sprint, and using agile methodologies in many more areas than software development. However, this is where the market is right now. The issue at hand is not so much the better application of business agility in the future, as applying today’s “agile BI” more effectively in the next 2-3 years. This is best done by recognizing one key fact: infusing the organization’s BI with agility is not primarily about development agility – it is about information agility.
Speeding up an inflexible business process via targeted analytics simply doubles your investment in and commitment to an information-handling process that needs instead the ability to fundamentally change. As I have noted in past posts, surveys show that in the typical process of receiving data and using analytics as part of the process of turning it into the right decision by the right decision-maker, or the right change in the business, the biggest problem is not that information is delivered too slowly. The fundamental problem is (a) that “leaks” at every step in the process result in more than 2/3 of the actionable information being lost along the way, and (b) the inflexibility of the process – specifically, its lack of openness to new types of data as they arrive – keep making the “leaks” worse.
Today’s best technology for dealing with (a) and (b) is something now called “data federation” or “data virtualization”. I won’t repeat the long litany of benefits to be expected from data virtualization; here, I will simply note that data virtualization, like master data management, provides one view of related data across the enterprise to the developer and the end user, and, unlike master data management, it can actively reach out, within or beyond the enterprise (as in Composite Discovery), to discover new data sources that can likewise be part of that global view of potential information. The one view of data tackles problem (a) at several critical points by ensuring that everyone can potentially see all data, and that analytics tools can be potentially applied to all useful data. The discovery feature tackles problem (b), if used effectively, by constantly refreshing the actionable data from, let’s say, new social-media data sources not presently covered by BI – as some Composite solutions do for Big Data.
Solving problem (b), of course, improves the organization’s information agility, not just its development agility. So the quick and effective method of tackling information agility in the here and now is to make a data virtualization development tool part and parcel of agile BI projects. Tools from folks like Composite Software, Tibco, Denodo, IBM, and SAP/Sybase are among those that would seem to fit the bill.
The Relevance of Composite Software to BI
Surprisingly enough, for a while around five years ago data virtualization’s predecessor, Enterprise Information Integration, was a hot topic in BI, as a way of extending data warehousing to operational data stores in order to query certain key new data in “near real time.” The primary results of this “fad” were to finally and definitively establish a major market for data virtualization, as well as to cement BI as Composite Software’s core market. In other words, unlike most or all other data-virtualization vendors, Composite Software has ten years of experience with real-world data virtualization at medium to large scales, five years of experience with BI at the same scales, a demonstrated ability to combine with other flexible BI tools such as MicroStrategy, and a focus on handling more “complex” needs that translates naturally to the more complex in-depth analytics typical of social-media and related Big Data analytics.
It may reasonably be objected that Composite, as well as other data virtualization vendors, has shown little or no interest in agile software development methodologies up to now, and therefore use of data virtualization solutions such as Composite’s may tend to slow down agile BI development efforts. Again, I would argue, as someone who has been touting agile development nearly since its inception twelve years ago, that this misses the point on two counts. In the first place, the rule of thumb with agile tools is “lead, follow, or get out of the way.” Where a SCRUM-supporting tool leads by urging developers towards agile best practices, and a really good refactoring tool follows in its wake by adding flexibility to the code thus produced, there is a third category of tools that simply automates and simplifies the necessary tasks of software development, while not getting in the way of the developer – and that is what something like Composite Information Server (CIS) with Composite Studio does for data-accessing functions in agile-BI programs.
In the second place, data virtualization such as Composite Software’s comes with a built-in gain in any agile-BI program’s information agility, which is often just as valuable as increased analytic-program implementation speed. Out of the box, the developer is assured of not missing a key in-enterprise data source. More subtly, the developer is made aware of other data sources than the one he or she may initially have targeted, which means that, as should be expected in agile development projects, new and unexpected features with value-add surface in the middle of the project.
In other words, while Composite Software may be in the middle of the pack when it comes to agility applied to agile-BI software development, it has long taken a leadership role in the features of data virtualization that really matter to agile-BI users, including BI services expertise, discovery, and complex data virtualization.
One final caveat appears to be going away. Since their inception, data virtualization tools have been oriented towards users above a certain size of medium-sized organization. Obviously, much of the action and excitement around agile BI has centered around its use in the public cloud by SMBs, who traditionally don’t deal with lots of various data sources that could use virtualization. However, what these SMBs are finding is that some public-cloud data is inherently “multi-organization”, and therefore requires federation and virtualization – especially since many of these same SMBs are now using multiple public clouds for their analytics. Sooner or later, Big Data has to be combined with ERP data on the cloud, and then with customer data on salesforce.com or wherever, and so on. As a result, even these SMBs are probably going to find data virtualization and Composite-type technology useful in their agile-BI efforts sooner rather than later.
Potential Uses of CIS-Type Agile BI for IT
Until recently, data virtualization in the enterprise has typically been done on a per-project basis, with the result that spreading it across all in-enterprise data is typically a work in progress. Obviously, this limits a little the usefulness of data virtualization in agile-BI development in the near future; but it is also an opportunity, in those enterprises that qualify, to kick-start the “global metadata repository” effort into high gear. Slap a data virtualization engine and development tool like CIS with Composite Studio into the agile-BI development team’s toolset, and have them scream to you about needing it to cover more data. That will get corporate’s attention, fast.
For those organizations that haven’t used data virtualization before, especially the SMBs, using data virtualization on things like Big Data is an excellent opportunity to avoid the arrows that the large-enterprise data-virtualization pioneer users faced. More specifically, if you start out by federating key customer Big Data with ERP on the cloud, private or public, you have already created a repository that covers most of your likely business-critical data in the next 2-3 years, and so you will not have to retrofit it later on.
A third, often underestimated, use of data virtualization is as a tool to support the new “business analysts” on the corporate side who are also doing agile analytics. A recent survey cited by Sloan Management Review suggests that these now view IT as slower and lower than dirt. Give them and support a tool that allows them to go out and discover new data sources, and I would not be surprised if they responded, like the woman conducting a choir in the movie Love Story, with “That is absolutely, stupendously, incredibly – OK.” That is, they won’t kiss you, but they will make a big shift towards tolerating you.
However, the use of CIS-type data virtualization that is dearest to my heart, and should be dearest to yours if you ever want to accomplish business agility rather than just BI “agility,” is applying its discovery features constantly, out on the Web, embedded in your analytics applications, to find and bring in-house new data and new data types. One key finding of business surveys is that only half of respondents find out about new key data heralding new market trends on the Web in less than half a year. Data virtualization can cut that to a day. You see, business agility isn't about reacting quickly and effectively to changes in your environment that you happen to find out about because they show up eventually in your customer complaints; the critical success factor is finding out all the changes that you can, as they happen out there, and then reacting quickly and effectively.
The Bottom Line for IT Buyers
Over the next few years, IT buyers should be viewing the acquisition and use of data virtualization tools in agile BI, among other things, as a no-brainer – but they don’t. Those smart IT buyers who do, however, will find that a pre-short-list is not necessary: on closer examination, the use cases almost present themselves. At any rate, that has been users’ experience in data virtualization’s original core markets such as the armed forces in the government – another Composite installed base.
However, for those constrained by budgets even in their strategic analytics initiatives, Composite and others may well have to go on a “pre-short list.” This should be of a peculiar type, however. The IT buyer should not passively wait for corporate to give IT a mandate that requires data virtualization. Instead, the IT buyer should be actively looking for the opportunity to introduce or extend agile-BI data virtualization, like the politician looking in every budget process or special spending initiative for a way of cutting taxes.
Composite Software is one among several who should be on that short list, as well as other data-virtualization-focused firms like Denodo and Tibco, not to mention the data federation solutions of IBM, Oracle via its BEA WebLogic Liquid Data acquisition, Red Hat/Metamatrix, and SAP/Sybase. Composite Software’s edge over most of these is simply that, over the last 10 years, Composite has established its experience and a leadership role in features and services useful to agile BI, and is among the most vocal about their plans to support Big Data and agile BI. Since that leadership role hasn’t diminished over the last three years, Composite Software will probably maintain its position at or near the top of agile-BI data virtualization short lists over the next 2-3 years, as well.
The bottom line for IT buyers of Other agile-BI data virtualization in general, and Composite Software’s CIS and Composite Studio in particular, is this: buy it right now or buy it a year from now, use it for BI agility or use it for business agility, but buy it and use it for some agile-BI initiative sometime in the next 2-3 years. Like an agile development methodology, it looks like it might cost more and cut into the bottom line, but it repeatedly costs less and delivers more revenue – because, used right, it improves your information agility.
The Importance of the Composite Software Technology to BI
Today’s enthusiasm for “agile BI” that often consists of slapdash applications of SCRUM to rapid roll-out of new analytics applets obscures the fact that business agility, not to mention New Product Development (NPD) agility, involves applying more of agile methodologies than a sprint, and using agile methodologies in many more areas than software development. However, this is where the market is right now. The issue at hand is not so much the better application of business agility in the future, as applying today’s “agile BI” more effectively in the next 2-3 years. This is best done by recognizing one key fact: infusing the organization’s BI with agility is not primarily about development agility – it is about information agility.
Speeding up an inflexible business process via targeted analytics simply doubles your investment in and commitment to an information-handling process that needs instead the ability to fundamentally change. As I have noted in past posts, surveys show that in the typical process of receiving data and using analytics as part of the process of turning it into the right decision by the right decision-maker, or the right change in the business, the biggest problem is not that information is delivered too slowly. The fundamental problem is (a) that “leaks” at every step in the process result in more than 2/3 of the actionable information being lost along the way, and (b) the inflexibility of the process – specifically, its lack of openness to new types of data as they arrive – keep making the “leaks” worse.
Today’s best technology for dealing with (a) and (b) is something now called “data federation” or “data virtualization”. I won’t repeat the long litany of benefits to be expected from data virtualization; here, I will simply note that data virtualization, like master data management, provides one view of related data across the enterprise to the developer and the end user, and, unlike master data management, it can actively reach out, within or beyond the enterprise (as in Composite Discovery), to discover new data sources that can likewise be part of that global view of potential information. The one view of data tackles problem (a) at several critical points by ensuring that everyone can potentially see all data, and that analytics tools can be potentially applied to all useful data. The discovery feature tackles problem (b), if used effectively, by constantly refreshing the actionable data from, let’s say, new social-media data sources not presently covered by BI – as some Composite solutions do for Big Data.
Solving problem (b), of course, improves the organization’s information agility, not just its development agility. So the quick and effective method of tackling information agility in the here and now is to make a data virtualization development tool part and parcel of agile BI projects. Tools from folks like Composite Software, Tibco, Denodo, IBM, and SAP/Sybase are among those that would seem to fit the bill.
The Relevance of Composite Software to BI
Surprisingly enough, for a while around five years ago data virtualization’s predecessor, Enterprise Information Integration, was a hot topic in BI, as a way of extending data warehousing to operational data stores in order to query certain key new data in “near real time.” The primary results of this “fad” were to finally and definitively establish a major market for data virtualization, as well as to cement BI as Composite Software’s core market. In other words, unlike most or all other data-virtualization vendors, Composite Software has ten years of experience with real-world data virtualization at medium to large scales, five years of experience with BI at the same scales, a demonstrated ability to combine with other flexible BI tools such as MicroStrategy, and a focus on handling more “complex” needs that translates naturally to the more complex in-depth analytics typical of social-media and related Big Data analytics.
It may reasonably be objected that Composite, as well as other data virtualization vendors, has shown little or no interest in agile software development methodologies up to now, and therefore use of data virtualization solutions such as Composite’s may tend to slow down agile BI development efforts. Again, I would argue, as someone who has been touting agile development nearly since its inception twelve years ago, that this misses the point on two counts. In the first place, the rule of thumb with agile tools is “lead, follow, or get out of the way.” Where a SCRUM-supporting tool leads by urging developers towards agile best practices, and a really good refactoring tool follows in its wake by adding flexibility to the code thus produced, there is a third category of tools that simply automates and simplifies the necessary tasks of software development, while not getting in the way of the developer – and that is what something like Composite Information Server (CIS) with Composite Studio does for data-accessing functions in agile-BI programs.
In the second place, data virtualization such as Composite Software’s comes with a built-in gain in any agile-BI program’s information agility, which is often just as valuable as increased analytic-program implementation speed. Out of the box, the developer is assured of not missing a key in-enterprise data source. More subtly, the developer is made aware of other data sources than the one he or she may initially have targeted, which means that, as should be expected in agile development projects, new and unexpected features with value-add surface in the middle of the project.
In other words, while Composite Software may be in the middle of the pack when it comes to agility applied to agile-BI software development, it has long taken a leadership role in the features of data virtualization that really matter to agile-BI users, including BI services expertise, discovery, and complex data virtualization.
One final caveat appears to be going away. Since their inception, data virtualization tools have been oriented towards users above a certain size of medium-sized organization. Obviously, much of the action and excitement around agile BI has centered around its use in the public cloud by SMBs, who traditionally don’t deal with lots of various data sources that could use virtualization. However, what these SMBs are finding is that some public-cloud data is inherently “multi-organization”, and therefore requires federation and virtualization – especially since many of these same SMBs are now using multiple public clouds for their analytics. Sooner or later, Big Data has to be combined with ERP data on the cloud, and then with customer data on salesforce.com or wherever, and so on. As a result, even these SMBs are probably going to find data virtualization and Composite-type technology useful in their agile-BI efforts sooner rather than later.
Potential Uses of CIS-Type Agile BI for IT
Until recently, data virtualization in the enterprise has typically been done on a per-project basis, with the result that spreading it across all in-enterprise data is typically a work in progress. Obviously, this limits a little the usefulness of data virtualization in agile-BI development in the near future; but it is also an opportunity, in those enterprises that qualify, to kick-start the “global metadata repository” effort into high gear. Slap a data virtualization engine and development tool like CIS with Composite Studio into the agile-BI development team’s toolset, and have them scream to you about needing it to cover more data. That will get corporate’s attention, fast.
For those organizations that haven’t used data virtualization before, especially the SMBs, using data virtualization on things like Big Data is an excellent opportunity to avoid the arrows that the large-enterprise data-virtualization pioneer users faced. More specifically, if you start out by federating key customer Big Data with ERP on the cloud, private or public, you have already created a repository that covers most of your likely business-critical data in the next 2-3 years, and so you will not have to retrofit it later on.
A third, often underestimated, use of data virtualization is as a tool to support the new “business analysts” on the corporate side who are also doing agile analytics. A recent survey cited by Sloan Management Review suggests that these now view IT as slower and lower than dirt. Give them and support a tool that allows them to go out and discover new data sources, and I would not be surprised if they responded, like the woman conducting a choir in the movie Love Story, with “That is absolutely, stupendously, incredibly – OK.” That is, they won’t kiss you, but they will make a big shift towards tolerating you.
However, the use of CIS-type data virtualization that is dearest to my heart, and should be dearest to yours if you ever want to accomplish business agility rather than just BI “agility,” is applying its discovery features constantly, out on the Web, embedded in your analytics applications, to find and bring in-house new data and new data types. One key finding of business surveys is that only half of respondents find out about new key data heralding new market trends on the Web in less than half a year. Data virtualization can cut that to a day. You see, business agility isn't about reacting quickly and effectively to changes in your environment that you happen to find out about because they show up eventually in your customer complaints; the critical success factor is finding out all the changes that you can, as they happen out there, and then reacting quickly and effectively.
The Bottom Line for IT Buyers
Over the next few years, IT buyers should be viewing the acquisition and use of data virtualization tools in agile BI, among other things, as a no-brainer – but they don’t. Those smart IT buyers who do, however, will find that a pre-short-list is not necessary: on closer examination, the use cases almost present themselves. At any rate, that has been users’ experience in data virtualization’s original core markets such as the armed forces in the government – another Composite installed base.
However, for those constrained by budgets even in their strategic analytics initiatives, Composite and others may well have to go on a “pre-short list.” This should be of a peculiar type, however. The IT buyer should not passively wait for corporate to give IT a mandate that requires data virtualization. Instead, the IT buyer should be actively looking for the opportunity to introduce or extend agile-BI data virtualization, like the politician looking in every budget process or special spending initiative for a way of cutting taxes.
Composite Software is one among several who should be on that short list, as well as other data-virtualization-focused firms like Denodo and Tibco, not to mention the data federation solutions of IBM, Oracle via its BEA WebLogic Liquid Data acquisition, Red Hat/Metamatrix, and SAP/Sybase. Composite Software’s edge over most of these is simply that, over the last 10 years, Composite has established its experience and a leadership role in features and services useful to agile BI, and is among the most vocal about their plans to support Big Data and agile BI. Since that leadership role hasn’t diminished over the last three years, Composite Software will probably maintain its position at or near the top of agile-BI data virtualization short lists over the next 2-3 years, as well.
The bottom line for IT buyers of Other agile-BI data virtualization in general, and Composite Software’s CIS and Composite Studio in particular, is this: buy it right now or buy it a year from now, use it for BI agility or use it for business agility, but buy it and use it for some agile-BI initiative sometime in the next 2-3 years. Like an agile development methodology, it looks like it might cost more and cut into the bottom line, but it repeatedly costs less and delivers more revenue – because, used right, it improves your information agility.
Thursday, January 5, 2012
The Other BI: EMC Greenplum and Embedded Analytics
This blog post highlights a software company and technology that I view as potentially useful to organizations investing in business intelligence (BI) and analytics in the next few years. Note that, in my opinion, this company and solution are not typically “top of the mind” when we talk about BI today.
The Importance of the Greenplum Software Technology to BI
I am stretching a point when I say that EMC’s Greenplum is not “top of the mind” today. EMC has done an extensive and effective job of marketing Greenplum’s virtues in dealing with Big Data. However, what I am talking about here is embedded analytics – and there, neither Greenplum nor any other vendor solution is “top of the mind” with IT today.
More specifically, I am talking about middle-tier analytics, the area that most embedded analytics will aim for in the next three years. This is not massive-data-store, in-depth-analytics BI like the data warehouse; nor is it the “smart sensor,” small-form-factor analytics that will increasingly come to the fore with the arrival of the sensor-driven Web (e.g., analytics on your iPhone). No, I am talking about medium-sized data stores, moderately in-depth analytics, and in-enterprise BI applied at the level of the department, local office, loosely-coupled storage array, or server network. This analytics does best when it is embedded in other software or in firmware, and operates semi-automatically to pick up business-process flows and alert the business before they get out of whack, or offloads load balancing from a central server. Unlike systems management software, embedded analytics not only monitors and “fixes” but also analyzes what is going on, and reports this analysis either to the top-tier data warehouse or a specific set of software, end users, and/or administrators.
Up to now, the fledgling beginnings of embedded analytics have begun to show up in the systems management software of folks like CA; but they are not separable pieces usable by other distributed software. Increasingly, the major vendors like IBM are now talking about taking analytic software from BI and analytics software suites and applying it to organization operations across the board.
However, these often involve databases retrofitted to BI in general and decision support in particular. What Greenplum represents is the obvious next step: applying a database designed from the ground up and optimized for querying and analytics. The point is that these will inevitably be better suited than data management approaches intended to handle updates as well as queries and result massaging.
This is not to say that an embedded analytics database is the end point of embedded-analytics evolution. Because most if not all available analytics databases were designed for the top tier, they are too “heavyweight” for their intended purpose: they perform more slowly, because they are tuned for much higher data-store sizes. However, whether the next turn of the market crowns a slimmed-down top-tier database or a new ground-up-designed middle-tier analytics database as the winner, either one will really do.
Over the next 2-3 years, it is reasonable for IT buyers to expect some of this technology to arrive on their doorsteps embedded in upgrades of existing solutions – but far from all of it. At some point in this period, separable analytics solutions will show up that will allow the user to go far beyond what a particular vendor is offering – if, of course, IT wants to.
Why would IT want to do this? Answer: to handle areas in which one-size-fits-all vendors are simply not moving fast enough. Take, for example, carbon accounting. Vendors have been very proactive in this area, but some of the market is moving faster still, towards monitoring that picks up on and alerts to excess emissions as they happen, and connects with the carbon accounting software when necessary. Likewise, as health care providers grapple with government mandates and Electronic Health Records, they can see coming a day in which they will need to perform damage control on breaches of privacy; but today’s tools are much slower than they could be to detect such a problem. In either case, customizable middle-tier embedded analytics that goes beyond most likely vendor offerings is needed.
The primary organization benefit of this technology, therefore, is deeper real-time understanding of in-enterprise problems that leads to better decision-making –a very cost-effective application of analytics’ general ability to improve gross margins. Embedded analytics via an analytics-adapted database may take longer to arrive than most of the Other BI that I talk about, but its advent and benefits are just as sure.
The Relevance of EMC to BI
While EMC has continued its tradition of “hands off the technology, add our markets” in the Greenplum acquisition, it has also continued another tradition: adding the technology where appropriate to its core storage software/firmware. That is, according to EMC (and I see no reason to doubt them), Greenplum technology is being put in storage controllers to offload querying from the server to the storage array. Obviously, that has a major positive implication for storage and large-BI performance. Less appreciated is the fact that this embedding of Greenplum requires that it “slim down” into a form that can operate not only on storage but also, in a middle-tier fashion, on loosely-coupled LANs serving local offices, departments, and so on. In other words, embedding on storage should mean that embedding on all other middle-tier form factors is within reach. And the acquisition of Greenplum also should mean that EMC is finally beginning to add database and BI smarts to its DNA, ensuring reasonable long-term service and support for its embedded-analytics solutions.
EMC’s market strength and apparent relative freedom from threat in the scale-out market mean that in the 2-3 year time frame I am talking about, and probably in the medium term as well, Greenplum is in no danger of going away. No, the real question for IT buyers of embedded analytics is whether EMC will have Greenplum take the next step, abstracting its slimmed-down form for embedded analytics on all vendor platforms. I can offer no guarantees of this, since it is not apparent that EMC has done such a thing before. All I can say is, if they do so, at least some sort of market will be there.
Potential Uses of Greenplum-Type Analytics for IT
It is time to point out that embedded-analytics technology is unusual in that vendors have relative freedom to delay delivering, say, multivendor or open-source middle-tier analytical databases, since it’s not high on IT wish lists. It could happen next week, or it could happen 3 years from now. So any IT acquisition of, and use of, this kind of embedded analytics will just have to wait until the vendors get around to it.
At that point, the obvious application is per-project – improving a specific business process or case-management implementation. More than other technologies, embedded analytics does not require full, integrated organizational implementation to be maximally effective. Rather, it does just fine applied to a task, a process, a function, a locality, or a local or strategic initiative. IT simply looks down the list of mission-critical projects and picks the one that benefits most from risk management or analytical automation.
The critical success factor in such projects is rapid implementation and upgrade, caused by automation of the implementation/upgrade process, allowing strategic projects a head start. Right now, while most vendors do well at this, high-end vendors like EMC seem to be setting the pace. And so, choosing EMC Greenplum (assuming it fits) in all likelihood means a better chance of rapid implementation and a database better fitted to a broad range of embedded-analytics tasks – not to mention better ongoing support for tricky cases.
The Bottom Line for IT Buyers
The IT buyer should view embedded analytics as a technology that may take a while to materialize. However, when it does, Greenplum-type embedded analytics will deliver analytics-type benefits at least equal to the whiz-bang high-end analytics now being sold – although those benefits will arrive in smaller per-project chunks. And that, in turn, means that this technology is definitely worth the IT buyer’s ongoing attention.
More specifically, the IT buyer might consider a “pre-pre-short-list” type of approach. That would involve identifying solutions such as EMC Greenplum that may wind up as part of the embedded-analytics short list in the next 2 years, and steadily moving those products in the pre-pre list over to the “pre-short list” as their technology reaches the point of usefulness (that is, it can be applied by IT rather than being embedded in another vendor solution, and it’s optimized for middle-tier analytics). Today, I would say that it appears Greenplum is probably among the closest to that take-off point. So, put it on the pre-pre short list, and get ready to put it on the short list. If everything goes right, and your CEO hits you with an urgent requirement that really demands embedded analytics, you will definitely be glad you had EMC’s Greenplum embedded analytics solution in your back pocket.
The Importance of the Greenplum Software Technology to BI
I am stretching a point when I say that EMC’s Greenplum is not “top of the mind” today. EMC has done an extensive and effective job of marketing Greenplum’s virtues in dealing with Big Data. However, what I am talking about here is embedded analytics – and there, neither Greenplum nor any other vendor solution is “top of the mind” with IT today.
More specifically, I am talking about middle-tier analytics, the area that most embedded analytics will aim for in the next three years. This is not massive-data-store, in-depth-analytics BI like the data warehouse; nor is it the “smart sensor,” small-form-factor analytics that will increasingly come to the fore with the arrival of the sensor-driven Web (e.g., analytics on your iPhone). No, I am talking about medium-sized data stores, moderately in-depth analytics, and in-enterprise BI applied at the level of the department, local office, loosely-coupled storage array, or server network. This analytics does best when it is embedded in other software or in firmware, and operates semi-automatically to pick up business-process flows and alert the business before they get out of whack, or offloads load balancing from a central server. Unlike systems management software, embedded analytics not only monitors and “fixes” but also analyzes what is going on, and reports this analysis either to the top-tier data warehouse or a specific set of software, end users, and/or administrators.
Up to now, the fledgling beginnings of embedded analytics have begun to show up in the systems management software of folks like CA; but they are not separable pieces usable by other distributed software. Increasingly, the major vendors like IBM are now talking about taking analytic software from BI and analytics software suites and applying it to organization operations across the board.
However, these often involve databases retrofitted to BI in general and decision support in particular. What Greenplum represents is the obvious next step: applying a database designed from the ground up and optimized for querying and analytics. The point is that these will inevitably be better suited than data management approaches intended to handle updates as well as queries and result massaging.
This is not to say that an embedded analytics database is the end point of embedded-analytics evolution. Because most if not all available analytics databases were designed for the top tier, they are too “heavyweight” for their intended purpose: they perform more slowly, because they are tuned for much higher data-store sizes. However, whether the next turn of the market crowns a slimmed-down top-tier database or a new ground-up-designed middle-tier analytics database as the winner, either one will really do.
Over the next 2-3 years, it is reasonable for IT buyers to expect some of this technology to arrive on their doorsteps embedded in upgrades of existing solutions – but far from all of it. At some point in this period, separable analytics solutions will show up that will allow the user to go far beyond what a particular vendor is offering – if, of course, IT wants to.
Why would IT want to do this? Answer: to handle areas in which one-size-fits-all vendors are simply not moving fast enough. Take, for example, carbon accounting. Vendors have been very proactive in this area, but some of the market is moving faster still, towards monitoring that picks up on and alerts to excess emissions as they happen, and connects with the carbon accounting software when necessary. Likewise, as health care providers grapple with government mandates and Electronic Health Records, they can see coming a day in which they will need to perform damage control on breaches of privacy; but today’s tools are much slower than they could be to detect such a problem. In either case, customizable middle-tier embedded analytics that goes beyond most likely vendor offerings is needed.
The primary organization benefit of this technology, therefore, is deeper real-time understanding of in-enterprise problems that leads to better decision-making –a very cost-effective application of analytics’ general ability to improve gross margins. Embedded analytics via an analytics-adapted database may take longer to arrive than most of the Other BI that I talk about, but its advent and benefits are just as sure.
The Relevance of EMC to BI
While EMC has continued its tradition of “hands off the technology, add our markets” in the Greenplum acquisition, it has also continued another tradition: adding the technology where appropriate to its core storage software/firmware. That is, according to EMC (and I see no reason to doubt them), Greenplum technology is being put in storage controllers to offload querying from the server to the storage array. Obviously, that has a major positive implication for storage and large-BI performance. Less appreciated is the fact that this embedding of Greenplum requires that it “slim down” into a form that can operate not only on storage but also, in a middle-tier fashion, on loosely-coupled LANs serving local offices, departments, and so on. In other words, embedding on storage should mean that embedding on all other middle-tier form factors is within reach. And the acquisition of Greenplum also should mean that EMC is finally beginning to add database and BI smarts to its DNA, ensuring reasonable long-term service and support for its embedded-analytics solutions.
EMC’s market strength and apparent relative freedom from threat in the scale-out market mean that in the 2-3 year time frame I am talking about, and probably in the medium term as well, Greenplum is in no danger of going away. No, the real question for IT buyers of embedded analytics is whether EMC will have Greenplum take the next step, abstracting its slimmed-down form for embedded analytics on all vendor platforms. I can offer no guarantees of this, since it is not apparent that EMC has done such a thing before. All I can say is, if they do so, at least some sort of market will be there.
Potential Uses of Greenplum-Type Analytics for IT
It is time to point out that embedded-analytics technology is unusual in that vendors have relative freedom to delay delivering, say, multivendor or open-source middle-tier analytical databases, since it’s not high on IT wish lists. It could happen next week, or it could happen 3 years from now. So any IT acquisition of, and use of, this kind of embedded analytics will just have to wait until the vendors get around to it.
At that point, the obvious application is per-project – improving a specific business process or case-management implementation. More than other technologies, embedded analytics does not require full, integrated organizational implementation to be maximally effective. Rather, it does just fine applied to a task, a process, a function, a locality, or a local or strategic initiative. IT simply looks down the list of mission-critical projects and picks the one that benefits most from risk management or analytical automation.
The critical success factor in such projects is rapid implementation and upgrade, caused by automation of the implementation/upgrade process, allowing strategic projects a head start. Right now, while most vendors do well at this, high-end vendors like EMC seem to be setting the pace. And so, choosing EMC Greenplum (assuming it fits) in all likelihood means a better chance of rapid implementation and a database better fitted to a broad range of embedded-analytics tasks – not to mention better ongoing support for tricky cases.
The Bottom Line for IT Buyers
The IT buyer should view embedded analytics as a technology that may take a while to materialize. However, when it does, Greenplum-type embedded analytics will deliver analytics-type benefits at least equal to the whiz-bang high-end analytics now being sold – although those benefits will arrive in smaller per-project chunks. And that, in turn, means that this technology is definitely worth the IT buyer’s ongoing attention.
More specifically, the IT buyer might consider a “pre-pre-short-list” type of approach. That would involve identifying solutions such as EMC Greenplum that may wind up as part of the embedded-analytics short list in the next 2 years, and steadily moving those products in the pre-pre list over to the “pre-short list” as their technology reaches the point of usefulness (that is, it can be applied by IT rather than being embedded in another vendor solution, and it’s optimized for middle-tier analytics). Today, I would say that it appears Greenplum is probably among the closest to that take-off point. So, put it on the pre-pre short list, and get ready to put it on the short list. If everything goes right, and your CEO hits you with an urgent requirement that really demands embedded analytics, you will definitely be glad you had EMC’s Greenplum embedded analytics solution in your back pocket.
Tuesday, January 3, 2012
The Other BI: Progress Apama and Event Processing
This blog post highlights a software company and technology that I view as potentially useful to organizations investing in business intelligence (BI) and analytics in the next few years. Note that, in my opinion, this company and solution are not typically “top of the mind” when we talk about BI today.
The Importance of the Apama Software Technology to BI
The value-add of Apama to BI, in my opinion, is the value-add of applying analytics to “data in motion” on a very broad range of data. Apama carries out “event processing”: conceptually, I think of event processing as a “processing head” monitoring an Enterprise Service Bus (ESB). Most data entering the organization, as well as data moving between data stores and between users within the organization, is wrapped up as messages and sent across the ESB to its destination. In the process, the “processing head” monitors the whole stream of data-in-motion and performs analytics, alerting, and other processing based on the type of data being reviewed (or, in the aggregate, the “pattern” of a stream of related data). What is unprecedented about this kind of data processing is that (a) it focuses on data across organizational units, unlike the typical data warehouse or operational database; (b) it intercepts some data the moment it arrives in the organization, which is the ultimate in real-time business-critical data processing; and (c) it can draw a direct line between that data and a decision-maker by alerting, so that business-critical decisions can be made as quickly as possible.
Practically, of course, an “event processor” can do (c) only for a certain small subset of the information in the organization, because by itself an event-processing database does not scale nearly as much as a data warehouse. The event processor has much less historical “context” as it processes each datum, because it simply does not have the time to perform a “query from hell” on terabytes of historical data before the next datum must be processed. In-depth analytics will simply have to wait, often for an hour or more. Nevertheless, this kind of instantaneous response is, in the real world, enormously valuable when fast response to the type of events that the event processor detects from the data is indeed mission-critical and/or business-critical.
At the same time, (a) – the ability to correlate data across organizational units – is an often-underestimated value of the event processor. As the discipline of systems analysis understands, a collection of business units is as much a set of process flows between units as a set of stand-alone companies. The job of corporate is often to ensure that these process flows work well, and the value-add of the event processor in this case is to provide enterprise performance management (EPM) that reads the tea leaves of particular process flows and ties them back to glitches in the performance of the units and their coordination. In plain English, a good event processor goes beyond what you could do before because it lets you respond immediately to some new threats and opportunities in your environment, and because it tells you some of the things that are really happening to muck up your business’ overall performance as it coordinates business units. If you combine the two, you get the famous “360-degree view” inside and outside the organization.
Apama, like other event processors, never operates in a vacuum. All organizations already have operational and decision-support databases supporting key applications, and event processing must adapt itself to handle what these do not. Therefore, Apama’s value-add within the limits of (a)-(c) above can vary quite widely. Always, however, if the user does a careful analysis of the most important decisions that need to be speeded up and the gaps in business performance information, the BI done by an event processor like Apama has a major impact on the organization – not on its bottom line, necessarily, but always on its business risk. The out-of-the-blue event that businesses always face becomes much less risky when an event processor manages to detect it in a timely fashion.
Right now, businesses of all sizes are still in the early stages of use of event processing – you can tell, because case studies typically trumpet particular per-project uses. Therefore, the field is wide open for approaches such as the “event-driven architecture”, in which an event processor on top of an ESB becomes the focal point at which corporate can not only monitor but also direct information flows; the “complex event processor”, in which on-the-fly analytics becomes far deeper; and “data streaming”, in which the whole notion of a data-warehouse database is upended to be a “processing head” handling querying on multiple parallel “streams” of XML-type data. These, however, will often not achieve full implementation until 2-4 years from now, at the least. The key value-add of event processing in real-world BI over the next 2-3 years, I believe, will be the ongoing identification and implementation of the most important alerts, decisions, and business-unit correlations that it can handle, and their integration with the existing BI architecture.
The Relevance of Progress Software to BI
The relevance of Progress Software itself to BI, and to data processing in general, is less clear than in its “glory days,” when (imho) it was a pioneer in near-lights-out database administration, rapid application development, Software as a Service (SaaS), and the ESB. In those days, it had a gift for identifying the innovative infrastructure-software simplification that led its SMB/departmental customers to rapid implementation of the latest large-enterprise functionality, and beyond. Apama arrived at the end of that period, as part of a successful Progress effort to provide services and software to allow its departmental customers to scale their SMB-type technology to the division, the line of business, and even the “edge” in the data center. Apama first found its niche in rapid analysis of massive streaming financial-market data; Progress Software appears, with its Event Manager, Event Modeler process-design end-user tool, SmartBlocks, and Dashboard Studio, to have added Progress’ own strengths in SMB-driven simplicity of use. To put it another way: Apama was born large-enterprise-ready; Progress added the veneer and tools that makes it fully in sync with open-source or agile BI as applied, say, to Big Data.
Thus, five years ago Progress Software would have been seen as an operational and decision-support database for an SMB, and a fast-moving local-level operational adjunct to enterprise BI in the Global 1000. Today, Progress Software has much less visibility in BI, and its connection to the latest BI technology is less visible; but if you look at the actual technology, Apama is indeed innovative and can deliver value-add across a wide range of enterprises. It only remains to ask, what’s the future of Progress Software as a supplier of event processing technology, and in general?
There are two parts to my answer. First, let me note the characteristic that Progress Software shares with just about every database company: it has a core loyal base of customers whose size may shrink, but whose tendency not to migrate away from the platform ensures that Progress Software will be extraordinarily long-lived. As in the past few years, database revenues may ebb over time; but few if any database companies in my 30 years of acquaintanceship with the industry see a massive collapse of their installed base. In the next 2-3 years, as sure as the sun rises, reasonable management plus this revenue flow will see Progress Software still standing (acquired or not) and still supporting Apama plus the infrastructure software like the Progress ESB that complements Apama.
However, there is little in the past three years, where revenues have been essentially flat, to suggest that Progress Software’s glory days will return, and that it will identify another new infrastructure technology to get strong growth started again. Moreover, Progress’ strength has never been that its solutions were a key part of the mainstream of Web innovation, and so there is no obvious reason to expect that Apama can, at a minimum, transition to form the core of cutting-edge open-source BI solutions. And that brings me to the second part of my answer.
I assert that these caveats almost certainly do not matter, in the next 2-3 years and probably further out. The reason is that Progress Software’s DNA may not be Web or open-source, but it is very definitely SMB-simple and flexible, with no vendor lock-in. Apama pre-acquisition would probably be a niche financial-market BI product. Apama plus Progress Software is a uniquely flexible and easy-to-use event processor that integrates with the rest of your architecture just fine, and it will stay that way. Progress Software’s new services prowess just ensures that the simplicity scales up to the largest of enterprises.
Potential Uses of DataRush for IT
The net of the contributions of both Apama technology and Progress Software’s “approach” to the Progress Apama solution is that, whether you are an SMB tackling BI for the first time or a large enterprise trying to become more agile in your querying of Big Data, Apama offers differentiated value-add, either stand-alone or as the complement to a large-vendor event processor and database architecture like IBM Streams and IBM Information Server. In the case of an SMB, the use case is straightforward: you can fit Apama with agile-BI efforts as a front end to catch urgent alerts implicit in Big Data, raw operational data, or ongoing analytics; and you can begin to develop EPM. Large enterprises typically have bottom-up Windows/desktop computing in parallel with the massive datacenter data warehouse: they can grow Apama with the evolution of that side of the enterprise architecture, while if appropriate also driving forward their top-down, datacenter-driven event processing projects using other vendors’ event-processing solutions, and easily integrating the two.
In either use case, the key to the most rapid possible success (I think) will be to identify on an ongoing basis the key targets for alerting and rapid decision-making, and to tie Apama as much as possible to historical data in order to allow the greatest “depth” of analytics at the point of data intercept. One good detection of a major customer about to dump you and immediate, effective reaction to prevent it will make the whole exercise more than worthwhile. And, remember, we are talking low-touch, very-low-TCO event processing here.
It is unusually hard to think of things that can go wrong with Progress Apama implementation. Services? Not as much needed, and Progress Software services that may be needed are already real-world-proven from more than 20 years of rave reviews from SMB and departmental clients, plus 5 years of driving similar technology into the division and LOB. Training? Again, the Progress Software track record is that even an untrained local-office manager can handle database maintenance – which is usually the biggest concern – and any developer can handle the drag-and-drop development tools. Integration? Progress Software has what I view as the standard set of adapters and gateways. Limits to scalability? Not if the Progress ESB is any guide. Let’s face it, IT implementation of Apama is very unlikely to be rocket science – and neither is gaining BI insights with it.
The Bottom Line for IT Buyers
At this point, an IT buyer should view Progress Software’s Apama as roughly equivalent to a “diamond in the rough.” It is the center of attention neither of BI, nor streaming technology, nor even sometimes of Progress Software itself. It suffers undeservedly from questions about Progress Software’s future, the future of event processing technology, whether Apama will continue to track BI technology, and a reputation as a high-end or specialized event processor. All it has going for it is that it is a more simple, more flexible, powerful event processing tool for a wide variety of use cases and scales, and should continue to deliver for the next 3-5 years, and almost certainly longer. And that should be plenty.
I have noted that Apama’s (and event processing’s) main value-add in BI is more in the risk area than in the top or bottom line (although some implementations, like EPM, do indeed impact revenues and costs). However, this is one technology that, when it succeeds, is really, really visible. Sell it to corporate as the latest technology fad if you like; the odds are that when the IT buyer acquires and then IT implements Apama, a big success story happens in the next year. And then you can concentrate on what’s really important: integration with the rest of your BI so that it all works optimally, in harmony.
The net-net for IT buyers, therefore, is to do a “reality check” on present-day event processing in BI, and then prepare a short list of event-processing software vendors to take the next step. I see no reason why, in most if not all cases, Progress Software’s Apama should not be on that Other BI “pre-short list.”
The Importance of the Apama Software Technology to BI
The value-add of Apama to BI, in my opinion, is the value-add of applying analytics to “data in motion” on a very broad range of data. Apama carries out “event processing”: conceptually, I think of event processing as a “processing head” monitoring an Enterprise Service Bus (ESB). Most data entering the organization, as well as data moving between data stores and between users within the organization, is wrapped up as messages and sent across the ESB to its destination. In the process, the “processing head” monitors the whole stream of data-in-motion and performs analytics, alerting, and other processing based on the type of data being reviewed (or, in the aggregate, the “pattern” of a stream of related data). What is unprecedented about this kind of data processing is that (a) it focuses on data across organizational units, unlike the typical data warehouse or operational database; (b) it intercepts some data the moment it arrives in the organization, which is the ultimate in real-time business-critical data processing; and (c) it can draw a direct line between that data and a decision-maker by alerting, so that business-critical decisions can be made as quickly as possible.
Practically, of course, an “event processor” can do (c) only for a certain small subset of the information in the organization, because by itself an event-processing database does not scale nearly as much as a data warehouse. The event processor has much less historical “context” as it processes each datum, because it simply does not have the time to perform a “query from hell” on terabytes of historical data before the next datum must be processed. In-depth analytics will simply have to wait, often for an hour or more. Nevertheless, this kind of instantaneous response is, in the real world, enormously valuable when fast response to the type of events that the event processor detects from the data is indeed mission-critical and/or business-critical.
At the same time, (a) – the ability to correlate data across organizational units – is an often-underestimated value of the event processor. As the discipline of systems analysis understands, a collection of business units is as much a set of process flows between units as a set of stand-alone companies. The job of corporate is often to ensure that these process flows work well, and the value-add of the event processor in this case is to provide enterprise performance management (EPM) that reads the tea leaves of particular process flows and ties them back to glitches in the performance of the units and their coordination. In plain English, a good event processor goes beyond what you could do before because it lets you respond immediately to some new threats and opportunities in your environment, and because it tells you some of the things that are really happening to muck up your business’ overall performance as it coordinates business units. If you combine the two, you get the famous “360-degree view” inside and outside the organization.
Apama, like other event processors, never operates in a vacuum. All organizations already have operational and decision-support databases supporting key applications, and event processing must adapt itself to handle what these do not. Therefore, Apama’s value-add within the limits of (a)-(c) above can vary quite widely. Always, however, if the user does a careful analysis of the most important decisions that need to be speeded up and the gaps in business performance information, the BI done by an event processor like Apama has a major impact on the organization – not on its bottom line, necessarily, but always on its business risk. The out-of-the-blue event that businesses always face becomes much less risky when an event processor manages to detect it in a timely fashion.
Right now, businesses of all sizes are still in the early stages of use of event processing – you can tell, because case studies typically trumpet particular per-project uses. Therefore, the field is wide open for approaches such as the “event-driven architecture”, in which an event processor on top of an ESB becomes the focal point at which corporate can not only monitor but also direct information flows; the “complex event processor”, in which on-the-fly analytics becomes far deeper; and “data streaming”, in which the whole notion of a data-warehouse database is upended to be a “processing head” handling querying on multiple parallel “streams” of XML-type data. These, however, will often not achieve full implementation until 2-4 years from now, at the least. The key value-add of event processing in real-world BI over the next 2-3 years, I believe, will be the ongoing identification and implementation of the most important alerts, decisions, and business-unit correlations that it can handle, and their integration with the existing BI architecture.
The Relevance of Progress Software to BI
The relevance of Progress Software itself to BI, and to data processing in general, is less clear than in its “glory days,” when (imho) it was a pioneer in near-lights-out database administration, rapid application development, Software as a Service (SaaS), and the ESB. In those days, it had a gift for identifying the innovative infrastructure-software simplification that led its SMB/departmental customers to rapid implementation of the latest large-enterprise functionality, and beyond. Apama arrived at the end of that period, as part of a successful Progress effort to provide services and software to allow its departmental customers to scale their SMB-type technology to the division, the line of business, and even the “edge” in the data center. Apama first found its niche in rapid analysis of massive streaming financial-market data; Progress Software appears, with its Event Manager, Event Modeler process-design end-user tool, SmartBlocks, and Dashboard Studio, to have added Progress’ own strengths in SMB-driven simplicity of use. To put it another way: Apama was born large-enterprise-ready; Progress added the veneer and tools that makes it fully in sync with open-source or agile BI as applied, say, to Big Data.
Thus, five years ago Progress Software would have been seen as an operational and decision-support database for an SMB, and a fast-moving local-level operational adjunct to enterprise BI in the Global 1000. Today, Progress Software has much less visibility in BI, and its connection to the latest BI technology is less visible; but if you look at the actual technology, Apama is indeed innovative and can deliver value-add across a wide range of enterprises. It only remains to ask, what’s the future of Progress Software as a supplier of event processing technology, and in general?
There are two parts to my answer. First, let me note the characteristic that Progress Software shares with just about every database company: it has a core loyal base of customers whose size may shrink, but whose tendency not to migrate away from the platform ensures that Progress Software will be extraordinarily long-lived. As in the past few years, database revenues may ebb over time; but few if any database companies in my 30 years of acquaintanceship with the industry see a massive collapse of their installed base. In the next 2-3 years, as sure as the sun rises, reasonable management plus this revenue flow will see Progress Software still standing (acquired or not) and still supporting Apama plus the infrastructure software like the Progress ESB that complements Apama.
However, there is little in the past three years, where revenues have been essentially flat, to suggest that Progress Software’s glory days will return, and that it will identify another new infrastructure technology to get strong growth started again. Moreover, Progress’ strength has never been that its solutions were a key part of the mainstream of Web innovation, and so there is no obvious reason to expect that Apama can, at a minimum, transition to form the core of cutting-edge open-source BI solutions. And that brings me to the second part of my answer.
I assert that these caveats almost certainly do not matter, in the next 2-3 years and probably further out. The reason is that Progress Software’s DNA may not be Web or open-source, but it is very definitely SMB-simple and flexible, with no vendor lock-in. Apama pre-acquisition would probably be a niche financial-market BI product. Apama plus Progress Software is a uniquely flexible and easy-to-use event processor that integrates with the rest of your architecture just fine, and it will stay that way. Progress Software’s new services prowess just ensures that the simplicity scales up to the largest of enterprises.
Potential Uses of DataRush for IT
The net of the contributions of both Apama technology and Progress Software’s “approach” to the Progress Apama solution is that, whether you are an SMB tackling BI for the first time or a large enterprise trying to become more agile in your querying of Big Data, Apama offers differentiated value-add, either stand-alone or as the complement to a large-vendor event processor and database architecture like IBM Streams and IBM Information Server. In the case of an SMB, the use case is straightforward: you can fit Apama with agile-BI efforts as a front end to catch urgent alerts implicit in Big Data, raw operational data, or ongoing analytics; and you can begin to develop EPM. Large enterprises typically have bottom-up Windows/desktop computing in parallel with the massive datacenter data warehouse: they can grow Apama with the evolution of that side of the enterprise architecture, while if appropriate also driving forward their top-down, datacenter-driven event processing projects using other vendors’ event-processing solutions, and easily integrating the two.
In either use case, the key to the most rapid possible success (I think) will be to identify on an ongoing basis the key targets for alerting and rapid decision-making, and to tie Apama as much as possible to historical data in order to allow the greatest “depth” of analytics at the point of data intercept. One good detection of a major customer about to dump you and immediate, effective reaction to prevent it will make the whole exercise more than worthwhile. And, remember, we are talking low-touch, very-low-TCO event processing here.
It is unusually hard to think of things that can go wrong with Progress Apama implementation. Services? Not as much needed, and Progress Software services that may be needed are already real-world-proven from more than 20 years of rave reviews from SMB and departmental clients, plus 5 years of driving similar technology into the division and LOB. Training? Again, the Progress Software track record is that even an untrained local-office manager can handle database maintenance – which is usually the biggest concern – and any developer can handle the drag-and-drop development tools. Integration? Progress Software has what I view as the standard set of adapters and gateways. Limits to scalability? Not if the Progress ESB is any guide. Let’s face it, IT implementation of Apama is very unlikely to be rocket science – and neither is gaining BI insights with it.
The Bottom Line for IT Buyers
At this point, an IT buyer should view Progress Software’s Apama as roughly equivalent to a “diamond in the rough.” It is the center of attention neither of BI, nor streaming technology, nor even sometimes of Progress Software itself. It suffers undeservedly from questions about Progress Software’s future, the future of event processing technology, whether Apama will continue to track BI technology, and a reputation as a high-end or specialized event processor. All it has going for it is that it is a more simple, more flexible, powerful event processing tool for a wide variety of use cases and scales, and should continue to deliver for the next 3-5 years, and almost certainly longer. And that should be plenty.
I have noted that Apama’s (and event processing’s) main value-add in BI is more in the risk area than in the top or bottom line (although some implementations, like EPM, do indeed impact revenues and costs). However, this is one technology that, when it succeeds, is really, really visible. Sell it to corporate as the latest technology fad if you like; the odds are that when the IT buyer acquires and then IT implements Apama, a big success story happens in the next year. And then you can concentrate on what’s really important: integration with the rest of your BI so that it all works optimally, in harmony.
The net-net for IT buyers, therefore, is to do a “reality check” on present-day event processing in BI, and then prepare a short list of event-processing software vendors to take the next step. I see no reason why, in most if not all cases, Progress Software’s Apama should not be on that Other BI “pre-short list.”
Labels:
analytics,
Apama,
BI,
event processor,
Progress Software,
SMB
Monday, January 2, 2012
The Other BI: Pervasive DataRush and Parallel Data Streaming
This blog post highlights a software company and technology that I view as potentially useful to organizations investing in business intelligence (BI) and analytics in the next few years. Note that, in my opinion, this company and solution are not yet typically “top of the mind” when we talk about BI today.
The Importance of the DataRush Software Technology to BI
The basic idea of DataRush, as I understand it, is to superimpose a “parallel dataflow” model on top of typical data management code, in order to improve the performance (and therefore scalability) of the data-processing operations used by typical large-scale applications. Right now, your processing in general and your BI querying in particular are typically done either by “query optimization” within a “database engine” that takes one stream of “basic” instructions and parallelizes it by figuring out (more or less) how to run each step in parallel on separate chunks of data, or by programmer code that attempts a wide array of strategies for speeding things up further, ranging from “delayed consistency” (in cases where lots of updates are also happening) to optimization for the special case of unstructured data (e.g., files consisting of videos or pictures). “Parallel dataflow” instead requires that particular types of querying/updates be separated into multiple streams depending on the type of operation. This is done up front, as a specification by a programmer of a dataflow “model” that applies across all applications with the same types of operation.
There is good reason to believe, as I do, that this approach can yield major, ongoing performance improvements in a wide variety of BI areas. In the first place, the approach should deliver performance improvements over and beyond existing engines and special-case solutions, and not force you into supporting yet another alternate technology path. The idea of dataflow is not new, but for various historical reasons this variant has not been the primary focus of today’s database engines, and so the job of retrofitting to support “parallel dataflow” is nowhere near completion in most database engines. That means that, potentially, using “parallel dataflow” on top of these engines can squeeze out additional parallelism, due to the increased number and sophistication of the streams, especially on massively parallel architectures such as today’s multicore-chip server farms.
At the same time, the increasing importance of unstructured and semi-structured data has created something of a “green field” in processing this data, especially in areas such as health care’s handling of CAT scans, vendors streaming video over the Web, and everyone querying social-media Big Data. Where existing data-processing techniques are not set in concrete, “parallel dataflow” is very likely to yield outsized performance gains when applied, because it operates at a greater level of abstraction than most database engines and special-case file handlers like Hadoop/MapReduce, and so can be customized more effectively to new data transaction mixes and data types.
There is always a caveat in dealing with “new” software technologies that are really an evolution of techniques whose time has come. In this case, the caveat concerns the fact that, as noted, programmers or system designers need to specify the dataflows, rather than the database engine, and this dataflow “model” is not a general case for all data processing. That, in turn, means that at least some programmers need to understand dataflows on an ongoing basis.
It is my guess that this is a task that users of “parallel dataflow” and DataRush should embrace. There is a direct analogy here between agile development and DataRush-based development. The usefulness of agile development lies not only in the immediate speedup of application development, but also in the way that agile development methodologies embed end-user knowledge in the development organization, with all sorts of positive follow-on effects on the organization as a whole. In the same way, setting up dataflows for a particular application leads typically to a new way of thinking about applications as dataflows, and that improves the quality and often the performance of every application that the organization handles, whether it is optimizable by “parallel dataflow” or not.
In other words, in my opinion, developers’ knowledge of data-driven programming is increasingly inadequate in many cases. Automating this programming in the database engine and user interface can only do so much to make up for the lack. It is more than worth the pain of additional ongoing dataflow programming to reintroduce the skill of programming based on a data “model” to today’s generation of developers.
The Relevance of Pervasive Software to BI
Let me state my conclusion up front: I view investment in Pervasive Software’s DataRush technology as every bit as safe as investment in an IBM or Oracle product. Why do I say this?
Let’s start with Pervasive Software’s “DNA.” Originally, more than 15 years ago, I ran across Pervasive Software as a spin-off of Novell’s Windows database of the 1980s. Over time, as databases almost always do, the solution that has become Pervasive PSQL has provided a stable source of ongoing revenue. More importantly, it has centered Pervasive Software from the very start in Windows, PC-server, and distributed database technologies servicing the SMB/large-enterprise-department market. In other words, Pervasive has demonstrated over 15 years of ups and downs that it is nowhere near failure, and that it knows the world even of the Windows/PC-server side of the Global 10,000 quite well.
At the same time, having followed the SMB/departmental market (and especially the database side) for more than 15 years, I am struck by the degree to which, now, software technologies move “bottom-up” from that market to the large enterprise market. Software as a Service, the cloud, and now some of the latest capabilities in self-service and agile BI are all taking their cue from SMB-style operations and technologies. Thus, in the Big Data market in particular and in data management in general, Pervasive is one leading-edge vendor well in tune with an overall movement of SMB-style open-source and other solutions centered around the cloud and Web data.
I therefore see the risks of Pervasive Software DataRush vendor lock-in and technology irrelevance over the next few years as minimal. And, of course, participation in the cloud open-source “movement” means crowd-sourced support as effective as IT’s existing open-source software product support.
Aren’t there any risks? Well, yes, in my opinion, there are the product risks of any technology, i.e., that technology will evolve to the point where “parallel dataflow” or its equivalent is better integrated into another company’s product. However, if that happens, dollars to doughnuts there will be a straightforward path from a DataRush dataflow model to that product’s data-processing engine – because the open-source market, at the very least, will provide it.
Potential Uses of DataRush for IT
The obvious immediate uses of DataRush in IT are, as Pervasive Software has pointed out, in Big Data querying and pharmaceutical-company grid searches. In the case of Big Data, DataRush front-ending Hadoop for both public and hybrid clouds is an interesting way to both reduce the number of instances of “eventual consistency” turning into “never consistent” and to increase the depth of analytics by allowing a greater amount of Big Data to be processed in a given length of time, either on-site at the social-media sites or in-house as part of handling the “fire hose” of just-arrived Big Data from the public cloud.
However, I don’t view these as the most important use cases for IT to keep an eye on. Ideally, IT could infuse the entire Windows/PC-server part of its enterprise architecture with “parallel dataflow” smarts, for a semi-automatic ongoing data-processing performance boost. Failing that, IT should target the Windows/small-server information handling in which increased depth of analytics of near-real-time data is of most importance – e.g., agile BI in general.
These suggestions come with the usual caveats. This technology is more likely than most to require initial experimentation by internal R&D types, and some programmer training, as well. Finding the initial project with the best immediate value-add is probably not going to be as straightforward as in some other cases, as the exact performance benefit of this technology for any kind of database architecture is apparently not yet fully predictable. Effectively, these caveats say: if you don’t have the IT depth or spare cash to experiment, just point the technology at a nagging BI problem and odds are very good that it’ll pay off – but it may not be a home run the first time out.
The Bottom Line for IT Buyers
Really, Pervasive DataRush is one among several performance-enhancing approaches that offer potential additional analytical power in the next few years, and so if IT passes this one up and opts for another, they may well keep pace with the majority of their peers. However, in an environment that most CEOs seem to agree is unusually uncertain, out-performing the majority, and extreme IT smarts in order to do so, are more frequently becoming necessary. At the least, therefore, IT buyers in medium-sized and large organizations should keep Pervasive DataRush ready to insert in appropriate short lists over the next two years. Preferably, they should also start the due diligence now.
The key to getting the maximum out of DataRush, I think, will be to do some hard thinking about how one’s BI and data-processing applications “group” into dataflow types. Pervasive Software, I am sure, can help, but you also need to customize for the particular characteristics of your industry and business. Doing that near the beginning will make extension of DataRush’s performance benefits to all kinds of existing applications far quicker, and thus will deliver far wider-spread analytical depth to your BI.
How will a solution like DataRush impact the organization’s bottom line? The same as any increase in the depth of real-time analysis – and right now that means that, over time, it will improve the bottom line substantially. For that reason, at the very least, Pervasive Software’s DataRush is an Other BI solution that is worth the IT buyer’s attention.
The Importance of the DataRush Software Technology to BI
The basic idea of DataRush, as I understand it, is to superimpose a “parallel dataflow” model on top of typical data management code, in order to improve the performance (and therefore scalability) of the data-processing operations used by typical large-scale applications. Right now, your processing in general and your BI querying in particular are typically done either by “query optimization” within a “database engine” that takes one stream of “basic” instructions and parallelizes it by figuring out (more or less) how to run each step in parallel on separate chunks of data, or by programmer code that attempts a wide array of strategies for speeding things up further, ranging from “delayed consistency” (in cases where lots of updates are also happening) to optimization for the special case of unstructured data (e.g., files consisting of videos or pictures). “Parallel dataflow” instead requires that particular types of querying/updates be separated into multiple streams depending on the type of operation. This is done up front, as a specification by a programmer of a dataflow “model” that applies across all applications with the same types of operation.
There is good reason to believe, as I do, that this approach can yield major, ongoing performance improvements in a wide variety of BI areas. In the first place, the approach should deliver performance improvements over and beyond existing engines and special-case solutions, and not force you into supporting yet another alternate technology path. The idea of dataflow is not new, but for various historical reasons this variant has not been the primary focus of today’s database engines, and so the job of retrofitting to support “parallel dataflow” is nowhere near completion in most database engines. That means that, potentially, using “parallel dataflow” on top of these engines can squeeze out additional parallelism, due to the increased number and sophistication of the streams, especially on massively parallel architectures such as today’s multicore-chip server farms.
At the same time, the increasing importance of unstructured and semi-structured data has created something of a “green field” in processing this data, especially in areas such as health care’s handling of CAT scans, vendors streaming video over the Web, and everyone querying social-media Big Data. Where existing data-processing techniques are not set in concrete, “parallel dataflow” is very likely to yield outsized performance gains when applied, because it operates at a greater level of abstraction than most database engines and special-case file handlers like Hadoop/MapReduce, and so can be customized more effectively to new data transaction mixes and data types.
There is always a caveat in dealing with “new” software technologies that are really an evolution of techniques whose time has come. In this case, the caveat concerns the fact that, as noted, programmers or system designers need to specify the dataflows, rather than the database engine, and this dataflow “model” is not a general case for all data processing. That, in turn, means that at least some programmers need to understand dataflows on an ongoing basis.
It is my guess that this is a task that users of “parallel dataflow” and DataRush should embrace. There is a direct analogy here between agile development and DataRush-based development. The usefulness of agile development lies not only in the immediate speedup of application development, but also in the way that agile development methodologies embed end-user knowledge in the development organization, with all sorts of positive follow-on effects on the organization as a whole. In the same way, setting up dataflows for a particular application leads typically to a new way of thinking about applications as dataflows, and that improves the quality and often the performance of every application that the organization handles, whether it is optimizable by “parallel dataflow” or not.
In other words, in my opinion, developers’ knowledge of data-driven programming is increasingly inadequate in many cases. Automating this programming in the database engine and user interface can only do so much to make up for the lack. It is more than worth the pain of additional ongoing dataflow programming to reintroduce the skill of programming based on a data “model” to today’s generation of developers.
The Relevance of Pervasive Software to BI
Let me state my conclusion up front: I view investment in Pervasive Software’s DataRush technology as every bit as safe as investment in an IBM or Oracle product. Why do I say this?
Let’s start with Pervasive Software’s “DNA.” Originally, more than 15 years ago, I ran across Pervasive Software as a spin-off of Novell’s Windows database of the 1980s. Over time, as databases almost always do, the solution that has become Pervasive PSQL has provided a stable source of ongoing revenue. More importantly, it has centered Pervasive Software from the very start in Windows, PC-server, and distributed database technologies servicing the SMB/large-enterprise-department market. In other words, Pervasive has demonstrated over 15 years of ups and downs that it is nowhere near failure, and that it knows the world even of the Windows/PC-server side of the Global 10,000 quite well.
At the same time, having followed the SMB/departmental market (and especially the database side) for more than 15 years, I am struck by the degree to which, now, software technologies move “bottom-up” from that market to the large enterprise market. Software as a Service, the cloud, and now some of the latest capabilities in self-service and agile BI are all taking their cue from SMB-style operations and technologies. Thus, in the Big Data market in particular and in data management in general, Pervasive is one leading-edge vendor well in tune with an overall movement of SMB-style open-source and other solutions centered around the cloud and Web data.
I therefore see the risks of Pervasive Software DataRush vendor lock-in and technology irrelevance over the next few years as minimal. And, of course, participation in the cloud open-source “movement” means crowd-sourced support as effective as IT’s existing open-source software product support.
Aren’t there any risks? Well, yes, in my opinion, there are the product risks of any technology, i.e., that technology will evolve to the point where “parallel dataflow” or its equivalent is better integrated into another company’s product. However, if that happens, dollars to doughnuts there will be a straightforward path from a DataRush dataflow model to that product’s data-processing engine – because the open-source market, at the very least, will provide it.
Potential Uses of DataRush for IT
The obvious immediate uses of DataRush in IT are, as Pervasive Software has pointed out, in Big Data querying and pharmaceutical-company grid searches. In the case of Big Data, DataRush front-ending Hadoop for both public and hybrid clouds is an interesting way to both reduce the number of instances of “eventual consistency” turning into “never consistent” and to increase the depth of analytics by allowing a greater amount of Big Data to be processed in a given length of time, either on-site at the social-media sites or in-house as part of handling the “fire hose” of just-arrived Big Data from the public cloud.
However, I don’t view these as the most important use cases for IT to keep an eye on. Ideally, IT could infuse the entire Windows/PC-server part of its enterprise architecture with “parallel dataflow” smarts, for a semi-automatic ongoing data-processing performance boost. Failing that, IT should target the Windows/small-server information handling in which increased depth of analytics of near-real-time data is of most importance – e.g., agile BI in general.
These suggestions come with the usual caveats. This technology is more likely than most to require initial experimentation by internal R&D types, and some programmer training, as well. Finding the initial project with the best immediate value-add is probably not going to be as straightforward as in some other cases, as the exact performance benefit of this technology for any kind of database architecture is apparently not yet fully predictable. Effectively, these caveats say: if you don’t have the IT depth or spare cash to experiment, just point the technology at a nagging BI problem and odds are very good that it’ll pay off – but it may not be a home run the first time out.
The Bottom Line for IT Buyers
Really, Pervasive DataRush is one among several performance-enhancing approaches that offer potential additional analytical power in the next few years, and so if IT passes this one up and opts for another, they may well keep pace with the majority of their peers. However, in an environment that most CEOs seem to agree is unusually uncertain, out-performing the majority, and extreme IT smarts in order to do so, are more frequently becoming necessary. At the least, therefore, IT buyers in medium-sized and large organizations should keep Pervasive DataRush ready to insert in appropriate short lists over the next two years. Preferably, they should also start the due diligence now.
The key to getting the maximum out of DataRush, I think, will be to do some hard thinking about how one’s BI and data-processing applications “group” into dataflow types. Pervasive Software, I am sure, can help, but you also need to customize for the particular characteristics of your industry and business. Doing that near the beginning will make extension of DataRush’s performance benefits to all kinds of existing applications far quicker, and thus will deliver far wider-spread analytical depth to your BI.
How will a solution like DataRush impact the organization’s bottom line? The same as any increase in the depth of real-time analysis – and right now that means that, over time, it will improve the bottom line substantially. For that reason, at the very least, Pervasive Software’s DataRush is an Other BI solution that is worth the IT buyer’s attention.
Wednesday, December 28, 2011
Methane Talk-Down: Partial
One of the true joys of learning about science – as opposed to, say, economics – is that eventually you can usually get to a scientific summary that clears up many of the distortions that popular reports create. In the midst of wading through yet another cherry-picked-evidence blog post (this one on methane) by Andrew Revkin of the NY Times, it suddenly occurred to me that I should check out Justin Gillis of the Times, whose posts have been praised iirc by Joe Romm of climateprogress fame. Gillis’ reporting still seemed a little superficial to me, but he had a link to a 2006 scientific summary of the research about methane and climate change, an oldie but goodie where I found the answers to many of my questions. My recent blog post on methane laid out the doomsday scenario that I fear; Chapter 6 of this summary, as Rachel Maddow would say, talked me down – but only partially.
Because the broad scenario that I laid out is not drastically affected by the information in the summary, it is easier to lay out the summary’s picture of methane and then, at the end, note how this may affect my scenario. I will focus on methane clathrates, since the changes to everything else are less substantial. And, of course, I am sure that more misconceptions remain – because a summary article of ongoing research can’t be expected to answer everything. Anyway, let’s begin.
Methane Clathrates, Water Methane
Last time, I presented a very summarized picture of natural-source methane as coming from three sources: methane clathrates under the sea, permafrost on land at high latitudes, and peat bogs next to the permafrost or in the tropics. It turns out that the picture is a bit more complicated, and the complications matter.
To start with, methane clathrates are formed and remain stable in sea-floor sediment in particular combinations of sea temperature and pressure from the sea above that limit them to the sea floor somewhere between 200 meters and 1000 meters below sea level. In other words, the water has to be near zero F, and the clathrate has to be lower than 200 meters below sea level and higher than 1000 meters below sea level. Between those two limits, the deeper the sea floor, the wider the zone in the sediment where it can exist. Guesstimates for a typical clathrate “stability zone depth” might be 250-300 meters. Btw, a confusing part of the scientific lingo apparently refers to Arctic clathrates as “subsea permafrost.”
What happens to melt the clathrates? The water next to the sea floor warms up, or warmer temps further up the sea slope cause the equivalent of a mudslide on the sea floor that basically slices through the clathrate, stirs up everything above the slice as a cloud of sediment, and melts all the clathrate above the new sea floor. That is what they think happened at Storegg, a place near Norway where there is a “crater” 30 km across that may have released a gigaton of carbon, all at once (methane is CH4).
Now here’s an odd part. We are used to thinking of gas coming up to the surface in bubbles and releasing itself into the atmosphere when the bubble pops. Not so with clathrate methane – most bubbles pop long before they rise the 200 meters or more to the surface, according to the models. Instead, one of several things happens: the methane rises to the surface but not as bubbles (it is “buoyant”) and then releases into the atmosphere, or it is eaten by methane-eating bacteria, or it converts (typically to carbon dioxide) en route. Initial indications are that a small percentage of melted clathrate should rise to the surface combined with water and is released into the atmosphere as methane, which happens effectively immediately; a large percentage should be eaten by bacteria, who convert it into carbon dioxide on the surface of the sea, and the carbon dioxide is released into the atmosphere in order to equalize atmospheric and oceanic CO2; and a medium-sized percentage should convert to carbon dioxide without going through the bacteria, to be released into the atmosphere as carbon dioxide in the same way.
The methane clathrates in the Arctic seas contain perhaps 50%-80% of all clathrates. They are also by far the most likely to be affected by global warming, since water temperature variation due to increased sunlight on the water and increased temps of sun-warmed currents from the south are widest there.
Other Methane Sources
The picture of land-based methane sources also needs amendment. It appears that much of the methane stored in permafrost is stored in peat within the permafrost – which can extend as far down as 200 meters or so. Meanwhile, wetlands at whatever latitude are generators of methane, the Amazon as much as Ireland. When the permafrost melts, the water plus peat turns into a bog that (under global warming) is maintained by increased precipitation: that’s what often drives increased methane production.
Here, the translation to the atmosphere is more clear. Melting of permafrost releases any methane locked in the ice (but not in clathrates), and also creates new constantly-emitting sources of methane. Likewise, wetlands inject methane directly into the atmosphere.
Now we come to the tricky part. We are accustomed to thinking of methane in the atmosphere as separate from carbon dioxide. Not so. What often happens to methane in the atmosphere is that it "oxidizes”, which typically means that one of the hydrogen atoms is broken off to help form H2O (water), while the rest forms a methyl group (CH3) which eventually breaks down to carbon dioxide. In other words, much of the methane tossed into the atmosphere actually winds up as the major greenhouse gas, and stays up there for 150-250 years.
What’s the Effect? Um …
OK, so now the scientist wants to figure out what the global-warming effect of unlocking all that methane is going to be. The problem is that we have two sources of comparison, and neither of them is great.
The first is to use what happens over 10-20,000 years immediately after a Milankovitch-cycle minimum (a “glaciation”) as a model. Using that model, scientists have pretty well determined that in such times of rising global temperatures, the amount of methane in the atmosphere probably doesn’t vary by a heck of a lot, and the effects on global temps compared to atmospheric carbon are pretty minimal. Methane melt in general might have a role in things like sea-ice melting near Greenland, which has been shown to have surprisingly wide effects on global climate, but most of the good candidates for that type of melt (subsea, permafrost, wetlands) just don’t make a strong case for themselves.
The problem with this type of analysis is that it looks only at periods when most of the ice remains – because that’s what happens at the peak temps of a Milankovitch cycle. We have almost certainly moved above those peak temps in the last couple of decades, and so we are in much less charted waters. For a period much more comparable, you have to go back to the PETM – 55 million years ago.
OK, in the PETM, temps were 5-10 degrees C warmer than now. Increases in carbon in the atmosphere just don’t seem to be enough to justify those warmer temps. So for a while, there were theories floating around that methane was the complete reason for that kind of warming – no carbon needed. That would have been nice, since figuring out why carbon suddenly spiked in the first place, not to mention why the time period of this rapid warming was around 20,000 years as the latest research suggests, has been a headache. Bad news: there simply doesn’t seem to be a natural source of methane that comes near to explaining the whole temperature rise, not to mention keeping going for 20,000 years. So it looks like we have a choice between carbon emissions plus “unknown”, and carbon plus methane. Tentatively, the scientists are voting for carbon plus methane.
But the PETM isn’t great as a model, either. The problem there is that things happened slowly compared to today. If we say that the carbon atmospheric-concentration rise then happened over the course of 20,000 years, well, our carbon rise appears to be happening over 350 years – and it may very well double the rise of the PETM over the course of those 350 years. In other words, this is happening at least a hundred times faster. And, as we’ve seen in the case of carbon, that can mean that the positive follow-on effects happen well before the negative “stabilizer” effects. So, for example, don’t necessarily expect the magical munching methane sea bacteria to appear in the Arctic and save the day.
OK, so the models we have aren’t great. Can we at least use them for some guesstimates?
Preliminary Guesses
Well, the scientists have done the guessing for me. The key sentences I find in Chapter 6 say, more or less (with the usual caveats about my understanding), that the amount of atmospheric methane from natural sources pre-Industrial Revolution equals the amount of methane added from human sources since then, which equals the likely amount of methane to be added at some point due to all natural sources except subsea methane, which equals the potential amount of methane from subsea methane. In other words, in a worst-case scenario with 2006 models, at some point in the next 300 years, we might expect atmospheric methane four times what it was in 1850.
How much added heating would that translate to? Again, reading between the lines, perhaps 1 degree C from the methane alone. However, if we take the PETM as a model, it might be more like 2 degrees C. And that’s the maximum, so we can all semi-relax, right?
Well, no. You see, there are two problems. First of all there’s the fact that much of that methane is going to convert to carbon dioxide when it’s up there. Second, there’s the fact that the more methane gets into the atmosphere from now on, the longer it sits there. The 2006 estimate was that methane hangs up in the atmosphere an average of 9 years. But at twice the concentration, I think we can count on it sitting up there for 12-18 years on average. So those two things should add another ½-1 degrees C to the “additive effect” of methane in the atmosphere.
And then, of course, there’s the question of methane that converts to carbon dioxide before it gets into the atmosphere. Here, the summary didn’t really have much to add in the long term. Even by their time-frame estimates, all that methane-to-carbon-dioxide, even if it doesn’t get there in the next 100 years, will almost certainly show up in the next 1000 years. So it’s a more extreme version of my “pay me now or pay me later” scenario – except that we can at least hope that by the time the methane-turned-carbon-dioxide shows up, we will have managed to cut down on our human-caused carbon emissions and the amount in the atmosphere will have begun to go downhill significantly.
All in all, not great, but not as bad as my full doomsday scenario. Instead of 6 degrees C from methane-turned-CO2, perhaps 2-3, although that increase will stick around for maybe twice as long; instead of 7-9 degrees C from methane-stayed-methane over the next 160 years, perhaps 1-2 degrees over the next 300-500 years. And it will happen more gradually, so it won’t be really noticeable, probably, for the next 30-40 years. Except …
I’m Not All the Way Down
Read carefully the interview with the head of the survey of methane releases in the Siberian Sea. He states, effectively, that the diameter of the “craters” I referred to earlier had increased by up to 100 times this year, and this methane was “bubbling to the surface.” If you look at the 2006 summary, neither is supposed to happen. Very little methane should arise to the surface in a bubble, as noted above, and the methane hydrates should not suddenly do a big jump in melting: and 5 degrees C increase in water temps (since 1984, a jump of 2.1 degrees C has been observed) should cause perhaps 1 meter’s worth or less of methane hydrates to melt over the next 40-80 years – and it can’t be explained as mudslides, since it has happened in quite a few places.
So why would scientists’ models be wrong? Well, in the first place, they assume that relative sea-water warming will only occur in a short space in the summer, when the ice is melted and the sun warms its top. However, the depth of the surface ice in winter is also less than before, and the water carried by currents from the south is warmer. Clearly, it’s very possible that scientists are underestimating the amount of melting going on the rest of the year. Add this to the known problems with the original model developed in 1995, and you have some, but maybe not all, of the increase in clathrate atmospheric methane release explained.
The second flaw may be the modeled prediction that very little methane melt will rise to the surface as bubbles. Why might this model be wrong? I don’t have a clear answer from the summary – it could be that the turbulence of the water keeps the bubble from popping, although that seems unlikely. One thing seems clear: the magical munching methane bacteria are nowhere to be seen.
And the third flaw, which also affects the land methane emissions rate, is a major underestimate in the models of the rate of global warming. The models implicitly assume that the Arctic sea ice won’t melt entirely in summer before somewhere between 2030 and 2100, and year-round perhaps never – that one seems clearly wrong. Therefore, they underestimate the speed of the follow-on effects, including much faster warming of water within 100 meters of the surface, which would inevitably mean much faster warming at the 200-500 meter level – sorry, that’s not “deep ocean.”
In other words, what the latest information is telling us is that the semi-comforting story I just gave you is almost certainly an underestimate. The “true” effect of methane is somewhere between my doomsday estimate and the one above – except that the roles of methane-stayed-methane and methane-turned-carbon-dioxide have switched, because we now know that much of that atmospheric methane is going to change to carbon dioxide while it’s up there.
I find the logic of the summary convincing as well as semi-comforting; so if I had to guess, I would say that the net effect is somewhere between 3-5 degrees C, mainly in carbon dioxide, and spiking over the next 40-150 years before leveling off. But that’s a complete guess. Until I understand just how the models went wrong, I’m only partially talked down from my panic. So here’s to the New Year: It will be a season of hope, it will be a season of despair, it will be a season of enormous impatience until the first scientific explanations come out.
Because the broad scenario that I laid out is not drastically affected by the information in the summary, it is easier to lay out the summary’s picture of methane and then, at the end, note how this may affect my scenario. I will focus on methane clathrates, since the changes to everything else are less substantial. And, of course, I am sure that more misconceptions remain – because a summary article of ongoing research can’t be expected to answer everything. Anyway, let’s begin.
Methane Clathrates, Water Methane
Last time, I presented a very summarized picture of natural-source methane as coming from three sources: methane clathrates under the sea, permafrost on land at high latitudes, and peat bogs next to the permafrost or in the tropics. It turns out that the picture is a bit more complicated, and the complications matter.
To start with, methane clathrates are formed and remain stable in sea-floor sediment in particular combinations of sea temperature and pressure from the sea above that limit them to the sea floor somewhere between 200 meters and 1000 meters below sea level. In other words, the water has to be near zero F, and the clathrate has to be lower than 200 meters below sea level and higher than 1000 meters below sea level. Between those two limits, the deeper the sea floor, the wider the zone in the sediment where it can exist. Guesstimates for a typical clathrate “stability zone depth” might be 250-300 meters. Btw, a confusing part of the scientific lingo apparently refers to Arctic clathrates as “subsea permafrost.”
What happens to melt the clathrates? The water next to the sea floor warms up, or warmer temps further up the sea slope cause the equivalent of a mudslide on the sea floor that basically slices through the clathrate, stirs up everything above the slice as a cloud of sediment, and melts all the clathrate above the new sea floor. That is what they think happened at Storegg, a place near Norway where there is a “crater” 30 km across that may have released a gigaton of carbon, all at once (methane is CH4).
Now here’s an odd part. We are used to thinking of gas coming up to the surface in bubbles and releasing itself into the atmosphere when the bubble pops. Not so with clathrate methane – most bubbles pop long before they rise the 200 meters or more to the surface, according to the models. Instead, one of several things happens: the methane rises to the surface but not as bubbles (it is “buoyant”) and then releases into the atmosphere, or it is eaten by methane-eating bacteria, or it converts (typically to carbon dioxide) en route. Initial indications are that a small percentage of melted clathrate should rise to the surface combined with water and is released into the atmosphere as methane, which happens effectively immediately; a large percentage should be eaten by bacteria, who convert it into carbon dioxide on the surface of the sea, and the carbon dioxide is released into the atmosphere in order to equalize atmospheric and oceanic CO2; and a medium-sized percentage should convert to carbon dioxide without going through the bacteria, to be released into the atmosphere as carbon dioxide in the same way.
The methane clathrates in the Arctic seas contain perhaps 50%-80% of all clathrates. They are also by far the most likely to be affected by global warming, since water temperature variation due to increased sunlight on the water and increased temps of sun-warmed currents from the south are widest there.
Other Methane Sources
The picture of land-based methane sources also needs amendment. It appears that much of the methane stored in permafrost is stored in peat within the permafrost – which can extend as far down as 200 meters or so. Meanwhile, wetlands at whatever latitude are generators of methane, the Amazon as much as Ireland. When the permafrost melts, the water plus peat turns into a bog that (under global warming) is maintained by increased precipitation: that’s what often drives increased methane production.
Here, the translation to the atmosphere is more clear. Melting of permafrost releases any methane locked in the ice (but not in clathrates), and also creates new constantly-emitting sources of methane. Likewise, wetlands inject methane directly into the atmosphere.
Now we come to the tricky part. We are accustomed to thinking of methane in the atmosphere as separate from carbon dioxide. Not so. What often happens to methane in the atmosphere is that it "oxidizes”, which typically means that one of the hydrogen atoms is broken off to help form H2O (water), while the rest forms a methyl group (CH3) which eventually breaks down to carbon dioxide. In other words, much of the methane tossed into the atmosphere actually winds up as the major greenhouse gas, and stays up there for 150-250 years.
What’s the Effect? Um …
OK, so now the scientist wants to figure out what the global-warming effect of unlocking all that methane is going to be. The problem is that we have two sources of comparison, and neither of them is great.
The first is to use what happens over 10-20,000 years immediately after a Milankovitch-cycle minimum (a “glaciation”) as a model. Using that model, scientists have pretty well determined that in such times of rising global temperatures, the amount of methane in the atmosphere probably doesn’t vary by a heck of a lot, and the effects on global temps compared to atmospheric carbon are pretty minimal. Methane melt in general might have a role in things like sea-ice melting near Greenland, which has been shown to have surprisingly wide effects on global climate, but most of the good candidates for that type of melt (subsea, permafrost, wetlands) just don’t make a strong case for themselves.
The problem with this type of analysis is that it looks only at periods when most of the ice remains – because that’s what happens at the peak temps of a Milankovitch cycle. We have almost certainly moved above those peak temps in the last couple of decades, and so we are in much less charted waters. For a period much more comparable, you have to go back to the PETM – 55 million years ago.
OK, in the PETM, temps were 5-10 degrees C warmer than now. Increases in carbon in the atmosphere just don’t seem to be enough to justify those warmer temps. So for a while, there were theories floating around that methane was the complete reason for that kind of warming – no carbon needed. That would have been nice, since figuring out why carbon suddenly spiked in the first place, not to mention why the time period of this rapid warming was around 20,000 years as the latest research suggests, has been a headache. Bad news: there simply doesn’t seem to be a natural source of methane that comes near to explaining the whole temperature rise, not to mention keeping going for 20,000 years. So it looks like we have a choice between carbon emissions plus “unknown”, and carbon plus methane. Tentatively, the scientists are voting for carbon plus methane.
But the PETM isn’t great as a model, either. The problem there is that things happened slowly compared to today. If we say that the carbon atmospheric-concentration rise then happened over the course of 20,000 years, well, our carbon rise appears to be happening over 350 years – and it may very well double the rise of the PETM over the course of those 350 years. In other words, this is happening at least a hundred times faster. And, as we’ve seen in the case of carbon, that can mean that the positive follow-on effects happen well before the negative “stabilizer” effects. So, for example, don’t necessarily expect the magical munching methane sea bacteria to appear in the Arctic and save the day.
OK, so the models we have aren’t great. Can we at least use them for some guesstimates?
Preliminary Guesses
Well, the scientists have done the guessing for me. The key sentences I find in Chapter 6 say, more or less (with the usual caveats about my understanding), that the amount of atmospheric methane from natural sources pre-Industrial Revolution equals the amount of methane added from human sources since then, which equals the likely amount of methane to be added at some point due to all natural sources except subsea methane, which equals the potential amount of methane from subsea methane. In other words, in a worst-case scenario with 2006 models, at some point in the next 300 years, we might expect atmospheric methane four times what it was in 1850.
How much added heating would that translate to? Again, reading between the lines, perhaps 1 degree C from the methane alone. However, if we take the PETM as a model, it might be more like 2 degrees C. And that’s the maximum, so we can all semi-relax, right?
Well, no. You see, there are two problems. First of all there’s the fact that much of that methane is going to convert to carbon dioxide when it’s up there. Second, there’s the fact that the more methane gets into the atmosphere from now on, the longer it sits there. The 2006 estimate was that methane hangs up in the atmosphere an average of 9 years. But at twice the concentration, I think we can count on it sitting up there for 12-18 years on average. So those two things should add another ½-1 degrees C to the “additive effect” of methane in the atmosphere.
And then, of course, there’s the question of methane that converts to carbon dioxide before it gets into the atmosphere. Here, the summary didn’t really have much to add in the long term. Even by their time-frame estimates, all that methane-to-carbon-dioxide, even if it doesn’t get there in the next 100 years, will almost certainly show up in the next 1000 years. So it’s a more extreme version of my “pay me now or pay me later” scenario – except that we can at least hope that by the time the methane-turned-carbon-dioxide shows up, we will have managed to cut down on our human-caused carbon emissions and the amount in the atmosphere will have begun to go downhill significantly.
All in all, not great, but not as bad as my full doomsday scenario. Instead of 6 degrees C from methane-turned-CO2, perhaps 2-3, although that increase will stick around for maybe twice as long; instead of 7-9 degrees C from methane-stayed-methane over the next 160 years, perhaps 1-2 degrees over the next 300-500 years. And it will happen more gradually, so it won’t be really noticeable, probably, for the next 30-40 years. Except …
I’m Not All the Way Down
Read carefully the interview with the head of the survey of methane releases in the Siberian Sea. He states, effectively, that the diameter of the “craters” I referred to earlier had increased by up to 100 times this year, and this methane was “bubbling to the surface.” If you look at the 2006 summary, neither is supposed to happen. Very little methane should arise to the surface in a bubble, as noted above, and the methane hydrates should not suddenly do a big jump in melting: and 5 degrees C increase in water temps (since 1984, a jump of 2.1 degrees C has been observed) should cause perhaps 1 meter’s worth or less of methane hydrates to melt over the next 40-80 years – and it can’t be explained as mudslides, since it has happened in quite a few places.
So why would scientists’ models be wrong? Well, in the first place, they assume that relative sea-water warming will only occur in a short space in the summer, when the ice is melted and the sun warms its top. However, the depth of the surface ice in winter is also less than before, and the water carried by currents from the south is warmer. Clearly, it’s very possible that scientists are underestimating the amount of melting going on the rest of the year. Add this to the known problems with the original model developed in 1995, and you have some, but maybe not all, of the increase in clathrate atmospheric methane release explained.
The second flaw may be the modeled prediction that very little methane melt will rise to the surface as bubbles. Why might this model be wrong? I don’t have a clear answer from the summary – it could be that the turbulence of the water keeps the bubble from popping, although that seems unlikely. One thing seems clear: the magical munching methane bacteria are nowhere to be seen.
And the third flaw, which also affects the land methane emissions rate, is a major underestimate in the models of the rate of global warming. The models implicitly assume that the Arctic sea ice won’t melt entirely in summer before somewhere between 2030 and 2100, and year-round perhaps never – that one seems clearly wrong. Therefore, they underestimate the speed of the follow-on effects, including much faster warming of water within 100 meters of the surface, which would inevitably mean much faster warming at the 200-500 meter level – sorry, that’s not “deep ocean.”
In other words, what the latest information is telling us is that the semi-comforting story I just gave you is almost certainly an underestimate. The “true” effect of methane is somewhere between my doomsday estimate and the one above – except that the roles of methane-stayed-methane and methane-turned-carbon-dioxide have switched, because we now know that much of that atmospheric methane is going to change to carbon dioxide while it’s up there.
I find the logic of the summary convincing as well as semi-comforting; so if I had to guess, I would say that the net effect is somewhere between 3-5 degrees C, mainly in carbon dioxide, and spiking over the next 40-150 years before leveling off. But that’s a complete guess. Until I understand just how the models went wrong, I’m only partially talked down from my panic. So here’s to the New Year: It will be a season of hope, it will be a season of despair, it will be a season of enormous impatience until the first scientific explanations come out.
Labels:
clathrates,
climate change,
global warming,
methane,
permafrost
Wednesday, December 21, 2011
Methane: The Final Shoe
Recently, Neven’s blog on Arctic sea ice (neven1.typepad.com) featured a new post on recent scientific observations of methane – observations that Neven said made him “sick to my stomach.” I am not as easily panicked – I reserved my stomach sickness for a recent British report about how most life in the oceans, except jellyfish, will be dead within the century unless we do far more than we are doing. However, I do understand his reaction. Effectively, these reports indicate that the dreaded “final shoe” of global warming, the one reinforcing side-effect of global warming that we hoped against hope would not happen, appears to be partially beginning to drop. Moreover, it seems clear to me that most if not all folks, even those who are aware of methane’s role in climate, are underestimating its potential impact in causing additional and more rapid murder, disaster, and then catastrophe.
So here is my understanding of methane’s role in our tragedy – for yes, some small tragedy is unavoidable now, even without methane’s impact and even if we do everything we can from now on. I am sure that as an amateur I am missing or misrepresenting some points. I am also pretty sure that most amateur commentators are doing far worse. If I were you, I would not take comfort from any of this post’s stumbles or missteps.
How It Works
Methane is CH4, or a carbon atom with four hydrogen atoms. It is, imho, the second most important “greenhouse gas”, and to understand its effects it is best to compare it with carbon tossed into the atmosphere and combined with oxygen to form CO2, or carbon dioxide.
Here’s how carbon emissions that form carbon dioxide work (excess detail stripped away). As they ascend in the air they combine with the oxygen to form carbon dioxide. Most of that carbon dioxide sits in the atmosphere for perhaps 150 years, and most of the carbon has fallen to earth again within 250 years. While it is up there, doubling the amount of carbon (in CO2 form) in the atmosphere adds about 3 degrees C to global temperatures, and double that in the far north and south, especially in the winter. The “normal” rate of carbon in the atmosphere is 250-280 ppm (parts per million), and we are presently somewhere around 395 ppm.
Methane emissions work in a similar fashion, but with some important differences. In the first place, methane is typically stored in the earth and emitted as a gas – i.e., not as carbon but as CH4. Once it gets into the air, it can either split the carbon atom to form carbon dioxide – hence increasing that greenhouse gas – or remain as methane. If it stays methane, most of it stays in the atmosphere for 10 years, and most is gone after 15 years. So what’s the problem?
Well, the problem is that while methane is in the atmosphere it has up to 70 times the impact on global warming of comparable amounts of carbon dioxide. I find it useful here to imagine the old image of keeping a ping-pong ball in the air with jets blowing from beneath. If I toss a ball of carbon dioxide in the air, it stays up there for 200 years. With methane, I have to keep blowing like crazy – or, if you like, adding the same number of new balls of methane every 12 years. But if I keep an amount of methane in the atmosphere comparable to doubling carbon dioxide, then I drive up temps not by 3 degrees C, but by 100 degrees C. No, we haven’t gotten to the worrisome part yet.
In other words, the effects of increased amounts of atmospheric methane, piled on top of increasing carbon dioxide from other sources, fall somewhere in between two extremes. At one extreme, all the methane turns into carbon dioxide, and hangs there for 150-250 years. As we will see, that means that carbon dioxide may double or quadruple compared to global warming without intervention by methane, for an additional 3-6 degree C global warming. At the other extreme, all the methane stays methane. As we will see, a reasonable guess for its effects then is a 12-18 degree C additional increase starting somewhere around 20 years from now and going for 120 years, and then fading out. To put it bluntly: we roast more now (stays methane) or we fry more later (changes to carbon dioxide).
The Real Worry
So where are these new methane emissions coming from? Mainly, there are two potential types of source. The first is human-caused activity: just as we emit more carbon by burning fossil fuels as our population and industry grows, so emissions of methane for industry and personal tasks and by increasing populations of farm animals like cows increases accordingly. The second is methane frozen in the earth between periods of unusual global warming. That methane lies in three main places:
i. The shallow Arctic seas, especially the shallow Siberian sea north of Russia, where it is frozen not in ice but in a comparable substance known as a clathrate;
ii. The permafrost of Siberia and northern Canada/Alaska, where it is locked in frozen ground tens of meters deep on top of unfrozen earth;
iii. The peat bogs further south, where it is often mixed with water.
The methane emissions from our first type of source have caused methane in the atmosphere to shoot up fairly steadily over the last 100-odd years, so that methane is now a significant contributor to today’s global warming. It would be a really excellent idea to cut down on it. However, to some extent that increase has apparently leveled off. No, what really scares us over the next 200 years is the second set of sources.
The last time the globe apparently became (not when it was, when it became) this warm or warmer – maybe 55 million years ago, in what is called the Paleocene-Eocene Thermal Maximum, or PETM for short – it seems very likely that methane from the second source type was indeed emitted in quantity as methane, as that is a very good explanation for why temps actually went a little higher than the amount of carbon dioxide in the air would seem to dictate. However, that should not give us comfort. Ken Caldeira in Nature notes that methane of this type was stored in much smaller quantities then. That would mean that methane from sources i-iii emitted now would either (a) have similar effects over a much longer period of time or (b) would have much greater effects over the same period of time. So which is it, (a) or (b)?
Well, one obvious factor in deciding between (a) and (b) is how fast our global temps are going up already, before we start emitting i-iii. Once we start that faster rate of emissions, of course, that will speed up global warming even further, so we can bet that a faster initial rate of global-temp increase will keep methane emissions higher right throughout the process. And every available bit of evidence points to the fact that we are warming already much faster than in the PETM – because we humans are emitting carbon stored in the earth as “fossil fuels” (really, mostly decayed vegetable matter) much faster.
All right, so it’s faster. Is it fast enough to worry about? Here we have to consider sources i, ii, and iii separately. Methane clathrates are apparently a big honking source of methane, according to scientific estimates. No one is entirely sure about how fast these clathrates will “melt” and methane bubbles will rise to the surface, once they start melting. However, they are sitting in shallow seas and they start right at the surface of the sea-floor. We know what it takes: warming of the water above the clathrates. And that has been happening, as the Arctic sea ice in that area at the top of the water melts in the summer where it hasn’t before, the sun beats down on the newly-exposed water to heat it, and warmer water from the south moves in.
Now let’s consider ii. Joe Romm at www.thinkprogress.com has an extended post focusing on this source. The net of what he has to say is: Methane stored in permafrost is comparable in amount to methane stored in clathrates – big and honking. Permafrost melting is already underway at a brisk pace. Projections that are unrealistically conservative about how fast global warming will occur project that methane/carbon dioxide emissions from that permafrost will reach a high level about 20 years from now and continue at that level or somewhat higher for 120 years, at which point most of the permafrost will be gone. Make your own adjustments – however you adjust, it’s going to reach a higher level than that sooner, and stay there for a shorter period of time. Is that enough, by itself, to worry about? You betcha.
Then there’s iii – peat bogs and wetlands, even in the tropics. It’s not clear that there is as much methane there, or that it will be released as quickly. Remember, the further south (north, for the Southern Hemisphere) you go, the slower the rate of global warming. But it’s very clear that it’s happening. That was what the Russian summer fires were all about: global warming led to warmer temperatures that dried up the peat bogs and they went up in smoke, releasing methane. My totally random guesstimate is that peat bog methane emissions will follow much the same trajectory as those in i and ii, and will therefore have ½ to ¼ the impact at any one time or overall of either i or ii.
Now let’s reach ahead and note that things get drastically worse if all of these emissions increases happen over the same period of time – somewhere around 2-2.5 times worse. Luckily, so far emissions source i has not yet kicked in. Scientists report that as of 2010, there were no atmospheric signs of unusual methane or CO2 from Arctic sea sources. Be careful. There’s a trick in that statement.
Doing the Math
Before we hurry on to our conclusion, let’s pause and see if we can nail down a little better what those effects are likely to be if all three sources of the second type fire off fast at the same time. In particular, let’s assume a scenario like that predicted for source ii, only this time with all three sources emitting like crazy.
One estimate has the amount of stored methane, converted to carbon dioxide, at about 7 times the amount in the atmosphere right now. Let’s assume that, starting 20 years from now, this emits about 1/3 of itself at a fairly steady rate over the course of the following 160 years. So, 2.3 times 400 ppm or 920 ppm is the amount added to the atmosphere by 2170, on top of the existing amount (400 ppm) and the amount estimated to be added to the atmosphere by 2100 under “business as usual” (900 ppm). We’re up to about 7-9 degrees C global warming somewhere between 2100 and 2150. Even if we cut our emissions to zero today (totally unrealistic), we’re up to 5-7 degrees Celsius.
If you want to be gloomy, you can assume almost all the methane, turned to carbon dioxide, vents in the same time period. Add on another 1800 ppm, and then add on another 500 ppm for continuing “business as usual” between 2100 and 2170. Now we’re talking 12 degrees C.
OK, same thing, but it all stays methane. If you remember, this is over 160 years, but methane falls from the sky after about 12 years, so we’re talking about 12/160 = 1/12.33 of the equivalent amount of carbon dioxide on an ongoing basis over that 160 years. However, that methane has perhaps 33 times the effect on temps while it’s up there. So starting 20 years from now, there is an overall jump of about 8 degrees C on top of the effect from carbon dioxide noted above – and that effect lasts for 160-odd years. That baked-in non-methane carbon dioxide effect is going to be around 3 degrees C under the most optimistic of assumptions, and could go as high as 9 degrees C. And remember, if you want to be gloomy, tack on an additional 6 degrees C from emitting almost all the methane.
So here’s your two extremes. If you’re lucky, it’ll all go up as methane, and fall right down again. Now we’re talking 11-23 degrees C global warming between 2030 and 2190, and we fall down to a nice comfortable 6 degrees C after that. If you’re unlucky, it’ll all go up as carbon dioxide, in which case we’ll see maybe 8 degrees C of global warming from 2050 to about 2330. By the way, initial estimates are that more of it will rise as carbon dioxide.
The best part of this analysis is that I left out other “positive forcings.” In particular, I left out the fact that all this warming is going to turn part of the ocean (The Arctic and Antarctic) and part of the land (all those nasty glaciers) darker, from white (snow) and off-white (ice) to dark brown and green (land) and dark blue (ocean). Darker colors store heat. I’m not sure what how much warming effect that will have, but scientific estimates suggest indirectly (it’s included in some scientific estimates suggesting 3-6 degrees additional warming beyond that due to carbon dioxide in the atmosphere) that it’s likely to be at least an additional 1 degree C. Icing on the cake. Or not icing.
The Best Part
Now back to that trick. You noticed, didn’t you, that I said “as of 2010.” Well, the recent article cited by Neven said that a Russian scientist reported that in their annual sample of Siberian Sea methane emissions, which they had been doing for 20 years, for the first time ever they were seeing, not “funnels” of tens of meters across from which methane was bubbling up, but lots of funnels “more than a thousand meters across”. Do the math: that’s between 1000 and 10,000 times the rate they had ever seen before. He was very confident that results across much of the Siberian Sea would be similar. Is that enough to signal the start of Arctic sea methane emissions on the scale we’ve been talking about? How can it not be enough?
Now let’s add the usual caveats: wide variance inherent in the estimates, lack of confirming evidence in some areas, uncertainties in data collection, blah, blah. The scientist’s reaction to these is to minimize the impact by stating the most likely impact of which he or she can be certain. The realistic reaction is to ask what is the impact of median likelihood, with equal likelihood of a lesser or greater impact – and, as far as I can tell, that’s what I’ve given you.
And, by the way, don’t bother to object that present projections don’t show this. Guess what – most models don’t consider the impact of even one natural methane source behaving this way, and the rest (only recently) of just one (permafrost).
Like I said, I don’t get sick to my stomach about this – because I did my own guesstimates more than a year ago and got my puking done then. I’m still hoping that the methane shoe will drop more slowly; and also that I’ll win the lottery. Right now, the latter seems more likely. Happy holidays, all.
So here is my understanding of methane’s role in our tragedy – for yes, some small tragedy is unavoidable now, even without methane’s impact and even if we do everything we can from now on. I am sure that as an amateur I am missing or misrepresenting some points. I am also pretty sure that most amateur commentators are doing far worse. If I were you, I would not take comfort from any of this post’s stumbles or missteps.
How It Works
Methane is CH4, or a carbon atom with four hydrogen atoms. It is, imho, the second most important “greenhouse gas”, and to understand its effects it is best to compare it with carbon tossed into the atmosphere and combined with oxygen to form CO2, or carbon dioxide.
Here’s how carbon emissions that form carbon dioxide work (excess detail stripped away). As they ascend in the air they combine with the oxygen to form carbon dioxide. Most of that carbon dioxide sits in the atmosphere for perhaps 150 years, and most of the carbon has fallen to earth again within 250 years. While it is up there, doubling the amount of carbon (in CO2 form) in the atmosphere adds about 3 degrees C to global temperatures, and double that in the far north and south, especially in the winter. The “normal” rate of carbon in the atmosphere is 250-280 ppm (parts per million), and we are presently somewhere around 395 ppm.
Methane emissions work in a similar fashion, but with some important differences. In the first place, methane is typically stored in the earth and emitted as a gas – i.e., not as carbon but as CH4. Once it gets into the air, it can either split the carbon atom to form carbon dioxide – hence increasing that greenhouse gas – or remain as methane. If it stays methane, most of it stays in the atmosphere for 10 years, and most is gone after 15 years. So what’s the problem?
Well, the problem is that while methane is in the atmosphere it has up to 70 times the impact on global warming of comparable amounts of carbon dioxide. I find it useful here to imagine the old image of keeping a ping-pong ball in the air with jets blowing from beneath. If I toss a ball of carbon dioxide in the air, it stays up there for 200 years. With methane, I have to keep blowing like crazy – or, if you like, adding the same number of new balls of methane every 12 years. But if I keep an amount of methane in the atmosphere comparable to doubling carbon dioxide, then I drive up temps not by 3 degrees C, but by 100 degrees C. No, we haven’t gotten to the worrisome part yet.
In other words, the effects of increased amounts of atmospheric methane, piled on top of increasing carbon dioxide from other sources, fall somewhere in between two extremes. At one extreme, all the methane turns into carbon dioxide, and hangs there for 150-250 years. As we will see, that means that carbon dioxide may double or quadruple compared to global warming without intervention by methane, for an additional 3-6 degree C global warming. At the other extreme, all the methane stays methane. As we will see, a reasonable guess for its effects then is a 12-18 degree C additional increase starting somewhere around 20 years from now and going for 120 years, and then fading out. To put it bluntly: we roast more now (stays methane) or we fry more later (changes to carbon dioxide).
The Real Worry
So where are these new methane emissions coming from? Mainly, there are two potential types of source. The first is human-caused activity: just as we emit more carbon by burning fossil fuels as our population and industry grows, so emissions of methane for industry and personal tasks and by increasing populations of farm animals like cows increases accordingly. The second is methane frozen in the earth between periods of unusual global warming. That methane lies in three main places:
i. The shallow Arctic seas, especially the shallow Siberian sea north of Russia, where it is frozen not in ice but in a comparable substance known as a clathrate;
ii. The permafrost of Siberia and northern Canada/Alaska, where it is locked in frozen ground tens of meters deep on top of unfrozen earth;
iii. The peat bogs further south, where it is often mixed with water.
The methane emissions from our first type of source have caused methane in the atmosphere to shoot up fairly steadily over the last 100-odd years, so that methane is now a significant contributor to today’s global warming. It would be a really excellent idea to cut down on it. However, to some extent that increase has apparently leveled off. No, what really scares us over the next 200 years is the second set of sources.
The last time the globe apparently became (not when it was, when it became) this warm or warmer – maybe 55 million years ago, in what is called the Paleocene-Eocene Thermal Maximum, or PETM for short – it seems very likely that methane from the second source type was indeed emitted in quantity as methane, as that is a very good explanation for why temps actually went a little higher than the amount of carbon dioxide in the air would seem to dictate. However, that should not give us comfort. Ken Caldeira in Nature notes that methane of this type was stored in much smaller quantities then. That would mean that methane from sources i-iii emitted now would either (a) have similar effects over a much longer period of time or (b) would have much greater effects over the same period of time. So which is it, (a) or (b)?
Well, one obvious factor in deciding between (a) and (b) is how fast our global temps are going up already, before we start emitting i-iii. Once we start that faster rate of emissions, of course, that will speed up global warming even further, so we can bet that a faster initial rate of global-temp increase will keep methane emissions higher right throughout the process. And every available bit of evidence points to the fact that we are warming already much faster than in the PETM – because we humans are emitting carbon stored in the earth as “fossil fuels” (really, mostly decayed vegetable matter) much faster.
All right, so it’s faster. Is it fast enough to worry about? Here we have to consider sources i, ii, and iii separately. Methane clathrates are apparently a big honking source of methane, according to scientific estimates. No one is entirely sure about how fast these clathrates will “melt” and methane bubbles will rise to the surface, once they start melting. However, they are sitting in shallow seas and they start right at the surface of the sea-floor. We know what it takes: warming of the water above the clathrates. And that has been happening, as the Arctic sea ice in that area at the top of the water melts in the summer where it hasn’t before, the sun beats down on the newly-exposed water to heat it, and warmer water from the south moves in.
Now let’s consider ii. Joe Romm at www.thinkprogress.com has an extended post focusing on this source. The net of what he has to say is: Methane stored in permafrost is comparable in amount to methane stored in clathrates – big and honking. Permafrost melting is already underway at a brisk pace. Projections that are unrealistically conservative about how fast global warming will occur project that methane/carbon dioxide emissions from that permafrost will reach a high level about 20 years from now and continue at that level or somewhat higher for 120 years, at which point most of the permafrost will be gone. Make your own adjustments – however you adjust, it’s going to reach a higher level than that sooner, and stay there for a shorter period of time. Is that enough, by itself, to worry about? You betcha.
Then there’s iii – peat bogs and wetlands, even in the tropics. It’s not clear that there is as much methane there, or that it will be released as quickly. Remember, the further south (north, for the Southern Hemisphere) you go, the slower the rate of global warming. But it’s very clear that it’s happening. That was what the Russian summer fires were all about: global warming led to warmer temperatures that dried up the peat bogs and they went up in smoke, releasing methane. My totally random guesstimate is that peat bog methane emissions will follow much the same trajectory as those in i and ii, and will therefore have ½ to ¼ the impact at any one time or overall of either i or ii.
Now let’s reach ahead and note that things get drastically worse if all of these emissions increases happen over the same period of time – somewhere around 2-2.5 times worse. Luckily, so far emissions source i has not yet kicked in. Scientists report that as of 2010, there were no atmospheric signs of unusual methane or CO2 from Arctic sea sources. Be careful. There’s a trick in that statement.
Doing the Math
Before we hurry on to our conclusion, let’s pause and see if we can nail down a little better what those effects are likely to be if all three sources of the second type fire off fast at the same time. In particular, let’s assume a scenario like that predicted for source ii, only this time with all three sources emitting like crazy.
One estimate has the amount of stored methane, converted to carbon dioxide, at about 7 times the amount in the atmosphere right now. Let’s assume that, starting 20 years from now, this emits about 1/3 of itself at a fairly steady rate over the course of the following 160 years. So, 2.3 times 400 ppm or 920 ppm is the amount added to the atmosphere by 2170, on top of the existing amount (400 ppm) and the amount estimated to be added to the atmosphere by 2100 under “business as usual” (900 ppm). We’re up to about 7-9 degrees C global warming somewhere between 2100 and 2150. Even if we cut our emissions to zero today (totally unrealistic), we’re up to 5-7 degrees Celsius.
If you want to be gloomy, you can assume almost all the methane, turned to carbon dioxide, vents in the same time period. Add on another 1800 ppm, and then add on another 500 ppm for continuing “business as usual” between 2100 and 2170. Now we’re talking 12 degrees C.
OK, same thing, but it all stays methane. If you remember, this is over 160 years, but methane falls from the sky after about 12 years, so we’re talking about 12/160 = 1/12.33 of the equivalent amount of carbon dioxide on an ongoing basis over that 160 years. However, that methane has perhaps 33 times the effect on temps while it’s up there. So starting 20 years from now, there is an overall jump of about 8 degrees C on top of the effect from carbon dioxide noted above – and that effect lasts for 160-odd years. That baked-in non-methane carbon dioxide effect is going to be around 3 degrees C under the most optimistic of assumptions, and could go as high as 9 degrees C. And remember, if you want to be gloomy, tack on an additional 6 degrees C from emitting almost all the methane.
So here’s your two extremes. If you’re lucky, it’ll all go up as methane, and fall right down again. Now we’re talking 11-23 degrees C global warming between 2030 and 2190, and we fall down to a nice comfortable 6 degrees C after that. If you’re unlucky, it’ll all go up as carbon dioxide, in which case we’ll see maybe 8 degrees C of global warming from 2050 to about 2330. By the way, initial estimates are that more of it will rise as carbon dioxide.
The best part of this analysis is that I left out other “positive forcings.” In particular, I left out the fact that all this warming is going to turn part of the ocean (The Arctic and Antarctic) and part of the land (all those nasty glaciers) darker, from white (snow) and off-white (ice) to dark brown and green (land) and dark blue (ocean). Darker colors store heat. I’m not sure what how much warming effect that will have, but scientific estimates suggest indirectly (it’s included in some scientific estimates suggesting 3-6 degrees additional warming beyond that due to carbon dioxide in the atmosphere) that it’s likely to be at least an additional 1 degree C. Icing on the cake. Or not icing.
The Best Part
Now back to that trick. You noticed, didn’t you, that I said “as of 2010.” Well, the recent article cited by Neven said that a Russian scientist reported that in their annual sample of Siberian Sea methane emissions, which they had been doing for 20 years, for the first time ever they were seeing, not “funnels” of tens of meters across from which methane was bubbling up, but lots of funnels “more than a thousand meters across”. Do the math: that’s between 1000 and 10,000 times the rate they had ever seen before. He was very confident that results across much of the Siberian Sea would be similar. Is that enough to signal the start of Arctic sea methane emissions on the scale we’ve been talking about? How can it not be enough?
Now let’s add the usual caveats: wide variance inherent in the estimates, lack of confirming evidence in some areas, uncertainties in data collection, blah, blah. The scientist’s reaction to these is to minimize the impact by stating the most likely impact of which he or she can be certain. The realistic reaction is to ask what is the impact of median likelihood, with equal likelihood of a lesser or greater impact – and, as far as I can tell, that’s what I’ve given you.
And, by the way, don’t bother to object that present projections don’t show this. Guess what – most models don’t consider the impact of even one natural methane source behaving this way, and the rest (only recently) of just one (permafrost).
Like I said, I don’t get sick to my stomach about this – because I did my own guesstimates more than a year ago and got my puking done then. I’m still hoping that the methane shoe will drop more slowly; and also that I’ll win the lottery. Right now, the latter seems more likely. Happy holidays, all.
Labels:
arctic sea ice,
carbon emissions,
global warming,
methane
Subscribe to:
Posts (Atom)