Thursday, 31 July 2014

Today I Learned...

Usenet-wannabe1 reddit has a section called "today I learned" in which people post links to various factoids that the presumably just discovered or otherwise  want to bring attention to.  It's a reasonable enough idea and the content is usually ok (aided by the potential for redditors to upvote & downvote posts that are worthy or misguided, respectively), if somewhat trended towards a black-and-white view of the world.

It's not my go-to place for learning interesting things, though, especially when there are shades of grey - or possibly ideas that aren't necessarily popular with either conservatives or (social) liberals. For me, that has to be the RSA, where some of the most interesting thinkers in the world come to put forward (admittedly usually along with a book) ideas that I'd either previously not considered or had wildly different views.

Another excellent source of thoughtful material are some of the science-y podcasts from Australia's ABC. In particular Future Tense and the Science Show go into depth and background for subjects that are all to often glossed over in the pursuit of a catchy headline or political opinion.

And that's the point: the interesting thing here is that in most cases, these are presentations by experts, not pundits. In an era where pretty much anyone can get their "point of view" on some kind of media2, it's refreshing to have a place to go to actually hear informed material - including informed and intelligent debate and (gasp) refutation - rather than the simplified version of most of the information we get elsewhere.

So - do yourself a favour, and get some of these podcasts (or watch the videos or read the transcripts) and be better informed as well as entertained.  I virtually never listen to the radio live these days.


[1] Ok - maybe not wannabe, but it's basically filling the role Usenet news held before all this new-fangled web business. #bah-humbug #whippersnappers

[2] Yes, I admit the irony in decrying the very thing I am doing...

Saturday, 26 July 2014

Cloud, high availability, antifragility and so on


In a previous article I wrote about how (the lack of) operational maturity may be impacting the adoption of private cloud in enterprise data centres.  In truth, that's really only half the story: the other significant question is "what applications can be run in the cloud" ?

The majority of "serious" enterprise applications have been around for a long time - think Oracle RDBMS, SAP CRM, and more.  Even if the full client-server or N-tier application stack of these systems has a distributed front-end, at their heart they typically run as monolithic programs that are very tightly integrated into their host computing system and its associated storage and other resources.

A great deal of modern IT design is focused on how to make these systems as resilient as possible, by deploying on robust underlying hardware and software infrastructure, providing redundancy within each host, and providing automated failover and disaster recovery systems to try to ensure there is no single point of failure. When more performance is needed, servers are upgraded with new capacity - hopefully during a seamless internal upgrade of CPU and or memory, but often via a carefully managed and often protracted "lift and shift" migration and update.

The thing is that to a great extent this concept of hardening, protecting, and updating a few known, vital systems runs counter to the "pure" cloud model, which is that cloud-based applications should be independent of the underlying platform & be able to simply scale up or down by adding instances for performance. The application itself should be "antifragile", that is, not need careful maintenance to ensure that it is up and running (this is the "pets vs cattle" analogy).

Antifragility is a term coined by Nassim Taleb (he of  the "Black Swan theory" fame) to describe something that does not merely withstand a shock but actually improves because of it.  Mr Taleb gives a great introduction to Antifragility in his speech at the RSA.  The idea is catching on in the industry: PWC describes how Instagram founders Mike Krieger and Kevin Systrom made use of the concept as they faced the immense problems of scaling their new platform.

The poster-children of cloud computing: Instagram, Netflix, and others, have built their applications (from the ground up) by adopting the antifragile approach. So far, however, traditional software vendors have not yet taken on this methodology, not least because the concepts are too radical for a large portion of their (understandably conservative) customer base.  In this respect we face a chicken-and-egg situation:  without a well-established base of private cloud computing environments to target, software application vendors are unlikely to create products with the cloud in mind.  Simultaneously, without applications that can take advantage of cloud infrastructure, operation managers and systems designers must rely on the traditional approach for making services highly-available, which these days can have the unfortunate side-effect of trying to shoehorn "pets" into an environment that's intended for "cattle"... which in turn yields poor results, frustration, and abandonment of "cloud" as an operational model.

As a "chicken-and-egg" problem there is an obvious solution[1]: start building the private cloud infrastructure for those applications that can make good use of it in the short term: development systems, stateless servers, short-term but frequently needed project infrastructure, and so on. Ideally, data centre managers can re-use their existing infrastructure and virtualisation systems under  Infrastructure-as-a-Service  (IaaS) platform management software, so as not to face a huge and complex migration or the additional expense of a separate silo of equipment just for "cloud".  Meanwhile many enterprise software vendors are working on Software-as-a-Services versions of their own products, in an attempt to capture that particular part of the market. This indicates that when and if cloud computing becomes a well-known operational method for private data centres, the software vendors have already done most of the work to "cloudify" their products.

The short version for IT managers and systems designers:  start building operational experience with private cloud now, and check with vendors about the availability of their products for "real" ("cattle-style") cloud deployment from time to time to assess the viability of moving mission-critical loads to a true cloud environment.



[1] In the case of "chicken-and-egg", the answer is "egg" (from dinosaurs, you see...). 

Saturday, 7 June 2014

Why Isn't Cloud Taking Off?


There's been some discussion in the press and from various pundits about how cloud infrastructure projects like OpenStack are "troubled", and how cloud hadn't taken off as expected, especially in the context of private clouds.

First of all, it's worth bearing in mind that in the quest for new news, the hype cycle of much of the media first tends to introduce something with great enthusiasm, make a lot of noise (and possibly overoptimistic predictions) about it, then when it fails to materialise according to the required new news schedule, announces a "failure" or otherwise turns against the original subject (you can see the same thing with reporting of celebrities).

So, cloud is still happening, and is probably inevitable even for the private cloud context at this stage, it's just not running to the timetable of noticeable results that the media would like to see.

But why is this? well it comes down to what "cloud" and especially "private cloud" really is.

At it's most fundamental, cloud computing is simply another way of handling operational management; that is, it's a methodology for organising your IT resources and making them available for use by your organisation.  Cloud computing is attractive because it promises to make more efficient use of resources (up to 80% utilisation, compared to approximately 50% for non-cloud virtualisation, or 30% for physical systems), and also because it promises more nimble access to those resources by project or business units via self-service. 

This efficiency and fast access relies on a number of factors, however. Firstly it relies on automation: without automation in systems configuration & service deployment, addition and access to resources is far too slow, and far to expensive. Cloud also relies on standardised templates for deploying services, even if those templates are as simple as "small", "medium", and "large" [1]. Without templates it's extremely difficult to apply rules to automatically assign workloads to the right systems, and also very hard to do proper pricing & chargeback. Thirdly, cloud relies on effective capacity management, so that new workloads don't suddenly choke the system - especially when end-users are able to self-serve their workloads.

If we look at these general needs of automation, templates, policy-based deployment, capacity planning and so on, we see that there needs to be a reasonably high level of operational maturity in an organisation to be successful at deploying cloud computing.

And this is where many organisations are struggling.

Research firm Gartner has an "operational maturity model" ranging from 0 ("running around with your hair on fire") through to 5 ("IT operations are integrated with the enterprise & provide additional value to the organisation").  The traits and mindset necessary to be truly successful with cloud computing exist around level 4, and perhaps an advanced stage of level 3.  Most IT organisations, for various reasons, usually operate in the region around level 2-3, with very few truly at level 4 and a vanishingly small number at level 5.


This is why, so far, private cloud computing deployments are few and far between: it is difficult stuff, and in many cases although organisations may truly want to adopt cloud operations, there remains a steep learning curve in many of the fundamental concepts that make cloud effective.

Organisations looking to implement cloud should look to develop their existing operational procedures, especially by making use of automation and tools that can handle templates, compliance,and capacity management.

Eventually cloud computing will become the standard method for operation, but there may need to be a generational change in organisational mindset, as well as in software & tools, before it can be fully realised.


Update:  this isn't the only reason, of course. Here's another thought.


1. Proper naming of service templates is important. Another article covers this.