Showing posts with label technology. Show all posts
Showing posts with label technology. Show all posts

Tuesday, 31 October 2017

How Open-Source Can Be the True Catalyst for Digital Change

[This article was first published in an abridged form in CIO Outlook magazine]

Twenty years ago, when the web was just starting to become available to consumers, a large proportion of non-tech websites was essentially advertising space: ephemeral virtual billboards promoting some aspect of an organisation’s product or service, but driven mostly by marketing departments with no solid connection to day to day business. Having been involved in developing website sin those days, and having often argued in vain with my customers that their web presence needed to be more than just a one-off disconnected experiment, it’s still remarkable to me that today the online world has so quickly become intrinsic to our lives. Today it is the disconnected organisation that is the anomaly: the first thing anyone does when starting up a new business these days is to check whether their proposed company name is available as a domain name, with changes made if it is not. Customers expect an online-first and mobile-first experience: we all know people who will skip past providers if they don’t have an easy-to-navigate site that lets them conduct business online, rather than having to telephone, or even worse: do something in person.

So, it’s a truism that customers expect to be able to access business services on their own terms, and to be able to do as much as possible online without having to resort to (potentially) slower methods of interaction. Businesses have to provide this ease of access or risk being passed by – but developing and maintaining the degree of interaction required is much more complex than the advertising websites of the 1990’s. In order to be truly effective, a company's digital entry-point must be able to reach directly into the workings of the organisation to be able to handle transactions in real-time and satisfy customer’s demands. This is the crux of digital transformation: the business processes that used to be kept internal to an organisation have to be codified and expressed in a way that makes it easy for customers to interact with the company… or else they will simply go to a competitor who can offer that experience.

This approach can be relatively easy when starting from scratch, but for established businesses the need to unravel years (or decades) of business logic and interconnected systems can become a nightmare, especially when time-to-market is important to satisfy customers’ ever-growing demands. Even for new companies, the need to constantly refresh and update quickly brings its own challenges: the market is never static, and if a competitor finds a new and more attractive way of providing the service, then a response needs to be found quickly.

Many organisations are turning to agile methodology and the associated concepts of DevOps and continuous integration to attempt to address the need for speedy time-to-market, yet the challenge for IT organisations is how to provide the systems needed to support the new approaches when up to 70% of their resources are spent just keeping the lights on in their day-to-day operations.

This is where open-source technology can help. In recent years the vast majority new systems innovation has arisen from the open-source arena; almost all cutting-edge technologies have an open-source aspect, with industry giants such as Google, Intel, IBM, and even Microsoft embracing open-source as a way to accelerate development and spread the adoption of the technologies underlying the modern web beyond the traditional proprietary boundaries. While open source software itself is not at all new, it has in recent years become the mainstream way of developing innovative ideas – based in most part on the foundation of that quintessential open source project: GNU/Linux. The freely-available nature of the Linux OS provides a platform with a level playing field for involvement, and the free GNU tools (compilers, libraries, and development environments) also lower entry barrier for new developers who can contribute towards community projects. Some of these projects can be tiny, perhaps with only a single contributor. Others, such as the OpenStack cloud infrastructure project, can include thousands of developers, project managers, and technical writers from big corporations, research organisations, and individuals.

The open nature of development means that more eyes and more ideas are brought to bear on a project: performance and security issues have a higher chance of being observed and resolved, with individual developers eager to make and maintain their personal credentials, and little to no opportunity for sweeping problems “under the carpet” to meet a specific deadline, as is the risk with closed, proprietary code. The open development also means that users aren’t locked-in to a particular vendor’s technology, or subjected to the risk of either unconscionable price hikes, or the prospect of a product being “killed” due to an acquisition or other business imperative. 

Of course, the challenge for IT managers when applying open source technology is how to support it – and especially how to support it whilst maintaining their existing systems. The widely-available nature of open source may make it very cost-effective to acquire, but those savings can be quickly eroded if an organisation has to employ their own experts to build, manage, maintain, and integrate those technologies, and can be a risky undertaking if those key employees leave the company for any reason (or even want to take a vacation). That’s where open-source software companies such as SUSE come in: over 25 years ago the Germany-based software company produced the first ever enterprise-ready version of Linux, and it has been building and integrating “infrastructure software” for enterprise ever since. This (profitable) longevity means that even though the technologies it products may be cutting-edge, the engineering and support is solid and reliable. A lot of this success is based on the hugely experienced development team, which makes up over 50% of the company’s employees. The egalitarian nature of open source development communities means that an individual developer’s personal credibility is extremely important in making a difference in the direction of an upstream project, and since SUSE boasts some of the most experienced developers in the industry, their influence can be seen across a wide range of projects, with the added result that the real-world scenarios they observe in customers’ workloads are considered when making changes or improvements.

The message, then, for CIO’s wanting help with digital transformation, is to look to the open source world for the reactive, adaptive, and innovative technologies that will make it possible to deliver on consumer expectations at a price point that is affordable, and to do so with the help of an experienced open-source partner who can provide the enterprise-grade support necessary to provide the stability needed for a reliable business.

Thursday, 7 July 2016

n-1 isn't necessarily the wisest choice

BMW 328 - original and hommage versions

Ask any vendor & you will find that one of their greatest frustrations is when customers insist on implementing only the "n-1" release of a particular product. 

At almost every meeting, vendors are asked about the availability of new features, or new capabilities, or new supported configurations that will match what the customer is trying to achieve, and yet when these are finally made available after much development and testing, customers will wait, and stick with supposedly safer older versions.

The risk-management logic, of course, is that the latest release is untried, and may contain flaws and bugs. Unfortunately this misses the point that fixes to older flaws are made possible by deploying a new release. It also brings up the laughable scenario of customers asking for new features to be "back-ported" to the older, "safe" release. Pro-tip: if you back-port all of your new features to the older release, then you end up with the new release anyway!

There are also some times when you just can't take advantage of the latest technology unless you're up-to-date: for example getting the most benefit out of new CPU's requires the operating system software to be in-sync. As SUSE VP of engineering Olaf Kirch points out in this article from 2012, when new features are introduced, you can either back-port to old code (possibly introducing errors) or take the new code and harden it. 

Which brings me to the real point of this article - when we're  talking about open source, the rate of change can be extremely rapid. This means that by the time you get a hardened, tested,  enterprise version of software out of the door, it is already at least version "n-1" : the bleeding-edge stuff is happening at the forefront of the community project, where many eyes and many egos are working on improvements to correctness and performance as well as features. So there's really no reason to require an n-1 release of, say, enterprise Linux ... all you're doing in that case is hobbling your hardware, paying more for extended support, and missing out on access to improvements.

So when SUSE introduces a new kernel revision mid-way through a major release, as it is doing with SUSE Linux Enterprise 12 Service Pack 2 (SLE12SP2), don't fret about the risks: the bleeding edge has already moved forward, and what you're getting is just the best, hardened, QA'd, engineered version of Linux with the most functionality.

Saturday, 2 April 2016

Why SUSE Linux Is the Only Sensible Choice for HPE Superdome-X


Unlike other material herein, this is an unashamedly partisan post. It's here's to collect together mostly links, for reference.

In December 2014, HP (now HPE) announced the successor to its long line of proprietary enterprise class computer systems:  the Integrity Superdome X.

What was particularly interesting about this announcement is that the focus was not just on the Intel Xeon processor, or the scalability of the machine to 16 CPU's and 24 TB of RAM, but what Jeff Kyle, director of product management for mission-critical systems at HP said was most important: "It's all about the software," Kyle told eWEEK.

Superdome X was the first flagship system from HP to not ship with HPUX.  Instead its launch OS was Linux. SUSE Linux.

Why?  Well SUSE Linux was the launch OS for Superdome X. And yes, that article mentions our old friends from Raleigh, but the fact remains that at launch time, the only benchmarks provided by HP were those with SUSE Linux.

So how about now, 18 months later?   Well according to the SPEC website, as of today ALL of the SpecJBB and SpecCPU benchmarks published by HPE for Superdome X use SUSE Linux Enterprise Server:



So there's a clear message here – when HPE is trying to get the best possible performance form their machines, they turn to SUSE Linux.

Why is this? well in terms of raw scalability, SUSE far exceeds RedHat linux:
Max supported CPUsMax supported RAM 
SLES 12:81921024 TB
RHEL 7.2:28812 TB


At the very least, this means that to handle a fully-loaded Superdome X with 24TB of RAM you must use SUSE Linux or risk falling into the "experimental" or at best "unsupported" category.  And risk isn't really want you want when running a mission-critical system.

In particular, SUSE Enterprise Linux has actually tested scalability on other systems to 8192 CPU cores and 64TB RAM (no-one could supply a machine with more RAM;  the CPU count was qualified after that article was written). SUSE's numbers are therefore not theoretical when it comes to the demands of Superdome X: there is no risk that scaling will not follow the expected path.

This difference is likely to continue as well: both HPE and Intel spend a lot of effort developing code for the Linux to improve scalability, etc, and this is sent “upstream” to the latest kernel versions.  This means that you would normally expect the most recent kernel version would have the best performance, due to HPE & Intel contributions.  Typically SUSE leads its competitor in implementing the latest kernel features whilst maintaining application and kernel binary compatibility between major kernel releases (e.g. from SLES 11SP1 with 2.6 kernel to SLES 11 SP2 with 3.0 kernel).  Given this background it should not be surprising for SUSE to continue to make similar advances in its current SLES 12 major release. Historically Red Hat has not changed major kernel releases:  RHEL 5 & 6 were kernel level 2.6, with 3.10 introduced only in 2014 with RHEL 7.

In other words: as HPE continues to contribute performance & other features to Linux for all of its server platforms, these are most likely to appear (with full global enterprise-class support) on SUSE Enterprise Linux long before they are available on RHEL. 



So it's no mistake that HP themselves chose SUSE Enterprise Linux when they migrated their internal systems from HP-UX.





Saturday, 22 August 2015

Why Open Source Software Defined Storage Will Win

Last week NetApp posted it's first quarterly loss in many years. It's not something I take pleasure in, since I have many friends still working at that company and it still has some very interesting technologies. The storm clouds are gathering though, and I can't help but liken NetApp's situation to that of Sun Microsystems, another former employer, as the tech bubble burst in the early 2000's, especially when I look at some of the comments reported from the NetApp earnings call.

Back in the day, Sun was making a lot of money and good margins with high performance proprietary hardware

Then along came Linux running some simple tasks on commodity hardware. it was good enough to do the simple jobs, had a lot of the essential features & functionality that unix provided, and the hardware was at such a low price that the economics were impossible to ignore. Some of the first adopters were the same companies that been the early sun customers, those who had replaced mainframes & minicomputers with cheaper, yet effective Sun hardware.

Unfortunately the lesson of their  own early success wasn't remembered by Sun's senior management, who thought they could win customers back by offering Linux on Sun hardware. Of course, the software wasn't the point here – it was the low-cost hardware that was attracting attention. Some of us in the field & parts of engineering tried to convince Sun's management to put more effort behind Solaris for x86, but just at a critical juncture, Solaris 9 was released on SPARC, along with the message that x86  support would be put on the back-burner ... the die was cast and Sun's fate sealed as commodity hardware pushed the boundaries of performance and drove the need for proprietary hardware into an upper niche. I still contend that if Sun had instead fully committed support for x86 & embraced a software-first approach,  the Solaris would dominate the market & Linux would have been relegated to a similar position as FreeBSD occupies today.

What does this have to do with software defined storage & NetApp? Well it appears that there is a similar approach from NetApp's senior management:  they see that customers "want scale out and software defined storage functionality",  but seem to think that the only response is a solution running on NetApp hardware. Like Sun, they (and the other major storage vendors) are chained to their high-margin proprietary hardware. Breaking free of this kind of entanglement is at the crux of Christensen's Innovator's Dilemma.

Meanwhile open source, software-defined storage solutions running on commodity hardware, such as SUSE Enterprise Storage based on Ceph – the Linux of enterprise storage if you like – are starting to gain attention, especially for large bulk data stores facing exponential growth and price-sensitivity. For the moment these solutions are best suited to relatively simple (albeit large) deployments, but the technology is not standing still. Ceph already features high-end technologies like snapshots, zero-copy cloning, cache-tiering and erasure coding, and enterprises are finding that it is "good enough" and at a dramatically lower price-point.  The open source nature of the development means that progress is rapid, and the software-defined nature means that hardware costs are driven relentlessly down. This is the same dynamics as we saw with the transition of UNIX to Linux, and it likely to have the same impact on the proprietary enterprise storage vendors.

So, change is here: I hope NetApp can navigate the course better than Sun did, though it will be interesting to see how.

Meanwhile, enterprises looking to rein in the exponential costs of enterprise storage can now look to open source for answers, and take advantage of the power and economics of commodity hardware systems.

Monday, 3 November 2014

Is Open Source Secure?


In the light of the recent Heartbleed and ShellShock vulnerabilities (both of which were apparently innocuous code bugs),  I was approached for comment for an article about the security of open source software.  The journalist (Anthony Caruana) wanted to know about how code gets approved for inclusion in open source projects, and especially about the potential for the inclusion of "malicious" code. This is a subject I touched on in a previous posting, so it was good to be able to address it more directly.

Here are the questions & my responses:

With such a huge community of developers contributed code to open source applications, what steps are in place to prevent malicious code being injected into widely distributed applications?

I take the term "malicious code"  to mean code that has been deliberately added for some nefarious purpose – to introduce a security back-door, or capture credentials and identifies, or do some other kind of damage. 
It's important to differentiate between this kind of deliberate attack, and the kind of programming errors that can result in exploitation, such as the openSSL Heartbleed bug.
The way that open source projects are managed actively works against the inclusion of malicious code. Firstly, the fact that the source is open and available to be audited by a large audience means that there is a much greater opportunity to find deliberate attacks than there is in closed source. It was source code audits that found the bugs in openSSL and Bash.  Malicious code would have to be carefully hidden - without leaving trace that it is being hidden – to avoid detection. Obfuscated code immediately attracts attention. 
The second factor is that major open source projects operate in a community where reputation is the critical factor:  simply submitting code is not necessarily enough to have it included into the main codeline.  The author must be known to, and trusted by,  the project maintainer.  It's human nature that when a new programmer joins a project, his or her code will be vetted more carefully than already-known contributors.  Establishing and maintaining a high reputation requires significant effort, so lends against the rapid or repeated insertion of malicious code. 
In contrast with this, closed-source code is not open for wide review by a large audience, and simply requires a disgruntled employee or the instruction of some agency to add the malicious code (as has been alleged by the Snowden leaks ).


What testing is done with open source software before it is released into the community? 

Testing varies from project to project - in this sense "open source" is not some homogenous group with identical policies. One of the features of the open source community is the rapid turn-around and release of projects:  with a major project there may be several different versions of software available at once - for example  a "stable" release, a "development" release, and a "nightly build".  A stable release can be considered to have gone through the most amount of testing; a development release is a version that developers and testers are working on; and a nightly build is version of the code that has the latest changes integrated & is almost certainly still buggy.  Of course, with commercial open source organisations like SUSE,  only stable versions are released.
Since this is open source it is the community itself that does the testing. The community may include commercial organisations like SUSE, but also includes researchers and hobbyists.  Users of open source software have to decide what level of comfort they have for possible bugs when they choose which version to use. One of the biggest contributions non-programmers can make is to try development release software & report back any problems.  Again - the potentially wider audience than you can achieve in a closed-source project means that this phase can be more effective and faster than in proprietary software (the "crowd source" effect). 
Hobbyists and non-programmers may only be able to investigate to a certain level, but commercial organisations and researchers can perform more involved tests. When SUSE contributes changes – either as new product features or patches – the software goes through our QA team to test functionality, stability, scalability and for regressions including vulnerabilities, as well as integration testing with our partners. SUSE makes heavy use of testing automation, and a lot of effort goes into maintaining & refining our automated methods. SUSE also relies on individual expertise to review results and find errors. 
At the end of the day for any software - open source, closed source or embedded - it's a question of confidence.  The attractive thing about open source is that there is potential for a much higher degree of confidence than for closed source.  It is much harder to hide mistakes, much harder to secretly introduce malicious code, and much more likely that a programmer wanting to make a name for him or herself will discover (and fix) problems when the source is open and available for review.  Ultimately, if the users of open source software want to perform rigorous and exhaustive examinations of code, they can;  the option is not even there for closed source software.

Saturday, 18 October 2014

Where IT Vendors lose customers

Having a mentor on hand to show the way can be vital for continued success.
Imagine your standard project deployment lifecycle:  3 months of research, 6 months of deployment, 3 years of operation, then it's time to start again.

The problem is, vendors tend to be involved at the "interesting" front-end of the cycle, but ignore the day-to-day side of things as they focus on the next win. This is hardly surprising: sales teams are continually goaled to make new revenue,  and are also mainly engaged with the big decision makers in an organisation, who typically are interested in the new projects rather than business as usual (BAU).

With three or more years of deployment, though, it's important to make sure that the new changes and techniques introduced in the project aren't lost, that once-winning features sidelined or relegated to the old way of doing things, or that – worst of all – some operational problem doesn't derail the entire system and throw the vendor's ability to execute into question.



As previously discussed, some of the sales-to-deployment transition can be better handled by incorporating solution architecture into the process: having someone shepherd the process from inception to deployment is important to make sure that everything goes to plan.

Even in cases where the deployment is carefully managed, however, many organisations still don't have a smooth handover from implementation to day-to-day operations.  This is the sort of case that devops is supposed to address – avoiding the "throw over the wall" mentality that is so common when IT projects move from the "interesting" new phase to the often-perceived boredom of normal operation.   In the worst cases, operations staff haven't been involved in the project, and are merely faced with a raft of new systems and procedures for which they don't see value.

This is where operations mentoring is important: getting a customer's BAU staff up to speed and onside with the new projects capabilities can be crucial to successful adoption of the technologies as intended, which in turn can have a significant impact on later sales.

The operational level is also a crucial relationship to make in order to understand what's going on at the customer, and avoiding "CNN moments" or unfortunate incidents just at the time that renewal (or another project) comes along.

An operations mentor is a skilled member of the vendor's technical staff – not necessarily part of the sales team – who can sit with the customer's operations group for an extended period (say three months), and help build up their own local knowledge in the new products and technologies. Where no local guru exists, this person fulfils that role, with the goal of bringing the customer's own team up to a higher level of operational capability.


It's important to realise that this isn't a traditional staff augmentation role: the goal here is to build up competency so that when the mentor leaves, the customer's team can continue on its own.  With that in mind, the tasks for the mentor should not be day-to-day management of the new systems, but should rather be to help the customer with supplementary education, tips and tricks, and the development of guides and run-books.  The idea is to follow the proverb: "give a man a fish and you feed him for a day; teach a man to fish and you feed him for a lifetime".

The "lifetime" in this context includes future interactions with the sales team – when a customer is comfortable with the previous technology, then they can be much more open to new products, and to deployment in new projects. Operations Mentoring is the way to develop that comfort, as well as to show that the customer's continuing success is important.

Saturday, 20 September 2014

Can Open Source Help Solve Unemployment?



The other day on the way from yet another airport to yet another hotel, I was chatting with the taxi driver who was interested in what I did. Inevitably, the taxi driver was not really a taxi driver, but just doing it as casual work while he looked for a real job.  He had a degree in electrical engineering, but like a lot of young people was finding it hard to get that first job since he lacked experience.

His story isn't unusual – according to The Smith Family's Dr Lisa O'Brien, Australia (like other countries) is facing record youth unemployment, with many candidates lacking job-ready skills.  Now having that university degree is probably going to help, but even this is no longer a guarantee of employment without experience.

So what has this got to do with Open Source?

Put simply, getting involved in an open source project is a great way for anyone to show that they can contribute in a meaningful way, work well with others, and develop skills and experience that can be directly transferred to a work environment.

The barrier to entry for open source projects is very low – you just need to show an interest in getting involved. In fact, it's not even necessary to be a proficient coder: open source projects often have more need of usability testers and documentation writers than programmers.  Although higher education can help with developing programming and project management skills, many open source projects have contributors who have not yet graduated or may not yet even be of university age.

The results can be quite dramatic: open source companies like SUSE frequently recruit new developers from the ranks of active contributors, and often look for open source experience & reputation rather than demanding formal qualifications.

What this means is that even without a particular degree or even paid work experience, involvement in open source can open doorways into an IT career, in a way that is relatively easy to access.

Just another reason why open source is increasingly important.



Wednesday, 3 September 2014

Bits and Bytes: SUSE® Cloud 4 OpenStack Admin Appliance – An Easier way to Start Your Cloud

Bits and Bytes: SUSE® Cloud 4 OpenStack Admin Appliance – An Easier way to Start Your Cloud: If you used the SUSE Cloud 3 OpenStack Admin Appliance, you know it was a downloadable, OpenStack Havana-based appliance, which even a non-technical user could get off the ground to deploy an OpenStack cloud.

Friday, 8 August 2014

Why is Open Source Important?



Recently I was asked by the IT manager of a customer, also a software development company, what my position was on open source. His developers were arguing in both directions: some characterised open source as being risky due to the potential for people to see the code & discover security vulnerabilities (which really isn't the case) ; other developers asserted that using closed source also had risks such as the vendor being acquired, going out of business, or dropping a particular product line.

My customer wanted to understand if had a "religion" about open vs closed, so I told a story about my perceptions on how closed source & closed systems came back to hurt the company that initiated it....

Microsoft defined the market in the late 1990's: most people couldn't see past the desktop paradigm. Certainly when I was at Sun Microsystems trying to tell people that "the Network is the Computer" I was mostly given blank stares or told I was living in a fool's paradise with the idea of continual network  access (now however,  some people even suffer from anxiety when they don't have network access).  Meanwhile, Microsoft, having been late to the start of the Internet revolution, quickly used its market power to inculcate Internet Explorer as the standard web browser for enterprise customers, so all software using a web interface had to confirm to its particular quirks, and in particular the quirks of Internet Explorer 6. Given this practical requirement, and the dominance of the Windows desktop concept, many software developers would not support their web interfaces with anything else, which in turn meant people had to buy into the Windows desktop world and so the vicious (or virtuous, depending on your point of view) circle continued.

When Microsoft tried to introduce new, better performing versions of Windows and Internet Explorer, however , they found that adoption was poor. Despite their best efforts and despite the end of support life for Windows XP (the last version of windows that supports internet explorer 6),  30% or more of Windows deployments are STILL of this older version, due in part to the huge dependance on the old, closed Microsoft ecosystem of the early 2000's.  In other words, Microsoft's own efforts to control the entire software stack actually ended up hurting them, with very poor adoption of Windows Vista, and initially slow adoptions of Windows 7 and 8.

A critical factor here is that these old systems do not work with today's dominant paradigm: mobile computing on phone or tablet.  Companies face a potentially huge transitional cost to get access to the way people now access data online. There are a lot of reasons why Microsoft hasn't been as dominant in the mobile space as they'd like to be, but certainly the way they set themselves up at the beginning of the 21st century didn't help.

So how does this fit into the question of open source?  Well by definition open source uses standards for data storage & transfer that are open to scrutiny and available for all to use. This means that getting to data and services can be possible from any device or system, not just from some tightly-coupled combination of software systems in a closed box owned by someone else.  Companies who developed with open standards in mind found it much easier (and faster) to move to the mobile world.

This is just one reason why open source is important: to provide tools, systems and protocols that can be continuously adapted & developed in compatible ways, rather than head down a one-way path towards a dead end.

A telling footnote to this story is that after many years of decrying open source, Microsoft is now embracing open source as a component of the way it operates, and actively collaborates with open source companies (such as SUSE) to improve interoperability and the effectiveness of their customers' IT environments.



Saturday, 26 July 2014

Cloud, high availability, antifragility and so on


In a previous article I wrote about how (the lack of) operational maturity may be impacting the adoption of private cloud in enterprise data centres.  In truth, that's really only half the story: the other significant question is "what applications can be run in the cloud" ?

The majority of "serious" enterprise applications have been around for a long time - think Oracle RDBMS, SAP CRM, and more.  Even if the full client-server or N-tier application stack of these systems has a distributed front-end, at their heart they typically run as monolithic programs that are very tightly integrated into their host computing system and its associated storage and other resources.

A great deal of modern IT design is focused on how to make these systems as resilient as possible, by deploying on robust underlying hardware and software infrastructure, providing redundancy within each host, and providing automated failover and disaster recovery systems to try to ensure there is no single point of failure. When more performance is needed, servers are upgraded with new capacity - hopefully during a seamless internal upgrade of CPU and or memory, but often via a carefully managed and often protracted "lift and shift" migration and update.

The thing is that to a great extent this concept of hardening, protecting, and updating a few known, vital systems runs counter to the "pure" cloud model, which is that cloud-based applications should be independent of the underlying platform & be able to simply scale up or down by adding instances for performance. The application itself should be "antifragile", that is, not need careful maintenance to ensure that it is up and running (this is the "pets vs cattle" analogy).

Antifragility is a term coined by Nassim Taleb (he of  the "Black Swan theory" fame) to describe something that does not merely withstand a shock but actually improves because of it.  Mr Taleb gives a great introduction to Antifragility in his speech at the RSA.  The idea is catching on in the industry: PWC describes how Instagram founders Mike Krieger and Kevin Systrom made use of the concept as they faced the immense problems of scaling their new platform.

The poster-children of cloud computing: Instagram, Netflix, and others, have built their applications (from the ground up) by adopting the antifragile approach. So far, however, traditional software vendors have not yet taken on this methodology, not least because the concepts are too radical for a large portion of their (understandably conservative) customer base.  In this respect we face a chicken-and-egg situation:  without a well-established base of private cloud computing environments to target, software application vendors are unlikely to create products with the cloud in mind.  Simultaneously, without applications that can take advantage of cloud infrastructure, operation managers and systems designers must rely on the traditional approach for making services highly-available, which these days can have the unfortunate side-effect of trying to shoehorn "pets" into an environment that's intended for "cattle"... which in turn yields poor results, frustration, and abandonment of "cloud" as an operational model.

As a "chicken-and-egg" problem there is an obvious solution[1]: start building the private cloud infrastructure for those applications that can make good use of it in the short term: development systems, stateless servers, short-term but frequently needed project infrastructure, and so on. Ideally, data centre managers can re-use their existing infrastructure and virtualisation systems under  Infrastructure-as-a-Service  (IaaS) platform management software, so as not to face a huge and complex migration or the additional expense of a separate silo of equipment just for "cloud".  Meanwhile many enterprise software vendors are working on Software-as-a-Services versions of their own products, in an attempt to capture that particular part of the market. This indicates that when and if cloud computing becomes a well-known operational method for private data centres, the software vendors have already done most of the work to "cloudify" their products.

The short version for IT managers and systems designers:  start building operational experience with private cloud now, and check with vendors about the availability of their products for "real" ("cattle-style") cloud deployment from time to time to assess the viability of moving mission-critical loads to a true cloud environment.



[1] In the case of "chicken-and-egg", the answer is "egg" (from dinosaurs, you see...). 

Saturday, 7 June 2014

Why Isn't Cloud Taking Off?


There's been some discussion in the press and from various pundits about how cloud infrastructure projects like OpenStack are "troubled", and how cloud hadn't taken off as expected, especially in the context of private clouds.

First of all, it's worth bearing in mind that in the quest for new news, the hype cycle of much of the media first tends to introduce something with great enthusiasm, make a lot of noise (and possibly overoptimistic predictions) about it, then when it fails to materialise according to the required new news schedule, announces a "failure" or otherwise turns against the original subject (you can see the same thing with reporting of celebrities).

So, cloud is still happening, and is probably inevitable even for the private cloud context at this stage, it's just not running to the timetable of noticeable results that the media would like to see.

But why is this? well it comes down to what "cloud" and especially "private cloud" really is.

At it's most fundamental, cloud computing is simply another way of handling operational management; that is, it's a methodology for organising your IT resources and making them available for use by your organisation.  Cloud computing is attractive because it promises to make more efficient use of resources (up to 80% utilisation, compared to approximately 50% for non-cloud virtualisation, or 30% for physical systems), and also because it promises more nimble access to those resources by project or business units via self-service. 

This efficiency and fast access relies on a number of factors, however. Firstly it relies on automation: without automation in systems configuration & service deployment, addition and access to resources is far too slow, and far to expensive. Cloud also relies on standardised templates for deploying services, even if those templates are as simple as "small", "medium", and "large" [1]. Without templates it's extremely difficult to apply rules to automatically assign workloads to the right systems, and also very hard to do proper pricing & chargeback. Thirdly, cloud relies on effective capacity management, so that new workloads don't suddenly choke the system - especially when end-users are able to self-serve their workloads.

If we look at these general needs of automation, templates, policy-based deployment, capacity planning and so on, we see that there needs to be a reasonably high level of operational maturity in an organisation to be successful at deploying cloud computing.

And this is where many organisations are struggling.

Research firm Gartner has an "operational maturity model" ranging from 0 ("running around with your hair on fire") through to 5 ("IT operations are integrated with the enterprise & provide additional value to the organisation").  The traits and mindset necessary to be truly successful with cloud computing exist around level 4, and perhaps an advanced stage of level 3.  Most IT organisations, for various reasons, usually operate in the region around level 2-3, with very few truly at level 4 and a vanishingly small number at level 5.


This is why, so far, private cloud computing deployments are few and far between: it is difficult stuff, and in many cases although organisations may truly want to adopt cloud operations, there remains a steep learning curve in many of the fundamental concepts that make cloud effective.

Organisations looking to implement cloud should look to develop their existing operational procedures, especially by making use of automation and tools that can handle templates, compliance,and capacity management.

Eventually cloud computing will become the standard method for operation, but there may need to be a generational change in organisational mindset, as well as in software & tools, before it can be fully realised.


Update:  this isn't the only reason, of course. Here's another thought.


1. Proper naming of service templates is important. Another article covers this.


Virtualization and Processor Capability


In the old days[1]  the big computer companies competed fiercely on processor performance: a few extra megahertz here & there could make the difference between winning that multimillion dollar sale & being relegated to the also-rans.  Applications and operating systems were monolithic, giant chunks of code that consumed vast quantities of the CPU's processing time, so the faster you could  run through that single-threaded code, the better.

These days, the "processor wars" are done & dusted. There's really nothing that the relatively low-volume-and-hence-high-price RISC processors found in the big UNIX systems from Sun, HP & IBM can do that the relatively high-volume-hence-low-price x86-64 processors can't. So unless you're looking for a specific niche solution (from Snoracle[2]) or especially clever I/O virtualisation (IBM Power), there's really not much reason to be putting your workloads on anything other than an x86-64  system.

In fact, processing capacity is so grotesquely huge these days (and has been for a while) that we chop that capacity up into little units & assign them to many operating systems at the same time. This is called "virtualisation" (stop me if I'm going too fast).   This leads to the question of whether we can get away with even less. Certainly in the dim dark ages of last century, we were able to run multiple web servers on a single cpu systems; even given the additional complexity of some web apps these days, there are many applications that don't warrant the full power possible even with virtualised x64, and certainly not the power & cooling requirements of those larger servers.

This is why people are interested in ARM and Atom processors - lower performance than the high-end x86-64, but with sufficient capacity to do the job, and much more attractive characteristics for power, cooling & price.

So the future of server selection could be much more about power consumption (& cooling costs) over its lifetime than raw performance.  In fact, it's almost certain.




1: Old days, for purposes of this article, are before about 2006. Bleh. I feel old.

2: Snoracle. Sun + Oracle. Also, possibly onomatopoeia .

Saturday, 7 December 2013

IP Longevity Through Open Source

You never forget working for your first vendor.  I was reminded of this at the SUSEcon conference last November, when I met a young colleague who had recently started with the company and was overflowing with enthusiasm for the culture, workmates, and feel of the workplace compared to his previous experiences.  I saw the same degree of enthusiasm in former colleagues at NetApp, who had grown up with the company, and invested themselves completely in the company culture, even as it went through change.  My own experience with Sun Microsystems was also a match: there was something special about working in a place which, especially for a technologist, offered access to such great ideas, people, equipment & opportunities. For me, this was compunded by being there during the dotcom boom.

While I was at Sun, the place seemed to be brimming with innovation - we were proud of our brilliant ideas & cutting-edge execution. There were a few less-enthusiastic, or perhaps better said as "more realistic" people who didn't look through such rose-coloured glasses. These were the folk who had come to Sun from elsewhere - places like DEC (Digital Equipment Company), which had previously been one of the go-to places & hotbeds of innovation. They could see the good times wouldn't last, and in the end were proved correct.


It seems there has been a constant stream of "innovation" companies, each in turn attracting enthusiastic contributors, building great technology, and then folding or being consumed by some larger organisation that ultimately fails to capitalise on the innovations. The tragedy here is that as the smart people leave these companies, and the intellectual property gets buried under legal constraints, the innovations get lost, and the next generation effectively has to start from scratch. Sun's amazing technology & concepts of the late 1990's & early 2000's has only recently become part of the mainstream understanding (in 1997, no-one could understand what "the network is the computer" meant - modern smartphones demonstrate the concept completely), yet a lot of companies today are re-inventing the capabilities of the last generation of technology Sun developed before the decline & loss of personnel.

This is where free & open source is so interesting: once ideas expressed in open source are exposed, they are forever available, so the demise of a particular company doesn't bury the intellectual property (even if it does mean a lot of the workers on a project may not be able to spend as much time on it). Perhaps as more and more development moves into open source, we'll have less re-invention of the wheel (or even less out & out ignorance of what has come before).  

So long as we can work out a way to agree on licensing schemes....



Wednesday, 3 July 2013

Open Cloud?

History has shown us how damaging it can be to funnel ourselves down a single proprietary path, no matter how popular a particular platform is at a particular point in time.  

It can even damage organisations who originally benefit from the lock-in: one only needs to look at the trouble Microsoft is having getting people to move on from Windows XP. A lot of that is due to the general move away from desktops to mobile platforms,  but early on a great deal was due to organisations having bought or built internal applications on top of Internet Explorer 6 & that did not work with Internet Explorer 7 or later, so would not work on newer Windows platforms....

Keep things open - make a bigger pie, rather than just trying to take all of the pie for yourself...

Friday, 18 January 2013

Going For Gold Isn't Enough


For many years, IT services within organisations have been advertised to business projects using a familiar set of names: usually “gold”, “silver” and “bronze”. There may be some noun associated with the service as well, such as “gold processing tier” or “silver storage tier”. The premise behind this kind of naming is that “everyone understands what this means”, and to some extent this is true from a relative point of view – gold is understood to be “better” than silver, which in turn is “better” than bronze. But what does “better” mean in the context of the services on offer? How can someone differentiate between the different services & decide which one to use? What happens when a new service is introduced that fits between the existing ones?

Most of the time the answer to these questions is “no-one knows”, which is less than ideal. The big problem with this precious-metal-related naming methodology (or any other arbitrary method such as colours, gemstones or suchlike), is that although they give a relative indication of “goodness”, they don't actually tell anyone what the service offers, and nor is there any consistency between what a name means from one organisation to the next (or even within a single organisation – they are arbitrary terms, after all). This usually results in a business project simply choosing what sounds like the “best” service (i.e. “gold”), even if the project's needs don't align with what the service is offering. Alternatively a cash-strapped project may choose the “least” option (i.e “bronze”) simply to save on costs, without understanding, for example, that running a database on bottom-tier infrastructure just won't work. This problem is compounded as IT organisations become more like service providers, and especially as automated orchestration is used more to provision and advertise services to business users through self-service portals and the like. It's this context which leads us to understand the way we should be designing and describing our IT services, which is to be as specific as possible.

The concept behind being a service provider (either internally to an organisation or externally), is to consolidate resources and then apportion them to different projects or customers on a per-use basis. The idea is that by offering a standard set of well-defined services, the provider can reduce waste, streamline processes, take advantage of infrastructure efficiencies like data deduplication and single-instance cloning, and save their organisation money (and/or make a margin). The key phrase here is “standard set of well-defined services”: in the service-provider context, customers (or projects) must choose from a menu, rather than being allowed to pick and choose how the ingredients will be combined. In turn, the service provider must be very clear on what is included in the services being offered, and also on what the price will be per unit.

In a NetApp storage context, this means that some understanding of potential workloads must exist in order to develop a service appropriate for each workload the provider wants to cater to, including all of the efficiencies and capabilities relevant to that workload. The goal is to then normalise the per-gigabyte pricing so that simple comparisons can be made on the effective usable space, rather than on the basic physical capacity. For example, likely de-duplication rates for Virtual Server (VSI) workloads in VMware are around 50%, which is greater than the de-duplication rates for simple file sharing at around 20%, so a provider can define different services with per-gigabyte rates for these two workloads that take these savings into account. For example, the “NFS datastore for VMware VSI” might be offered at $0.50/GB, while the “NFS datastore for general data” might be offered at $0.80/GB. This naming and pricing helps the customer or project understand what they are ordering, and helps the service provider direct the customer to the most appropriate service for a workload.

Once you start naming these services in a descriptive way, you can extend them to add additional capabilities. For example the “NFS datastore for VMware VSI” might be extended to a new service called “NFS datastore for VMware VSI with local backup”, which introduces NetApp SnapShots on a defined schedule, and includes an additional per-gigabyte price to cover to cost of capacity used for those backups (perhaps a 20% premium). Another example might be “NFS datastore for general data with Disaster Recovery”, which would include a SnapMirror copy of the data at a remote site, and would have an associated premium to cover both copies of the data, networking costs, and so on.

It's easy to see how infrastructure policies can be applied to basic services to build up a complete catalogue of services, and how the capabilities and efficiencies of the infrastructure can be used to set pricing so that customers or projects can select based on both functional requirements and budget, and service providers can direct customers to the most appropriate infrastructure. With a descriptive naming system, providers can easily introduce new or extended services as technologies become available, without worrying about having to squeeze between or around arbitrary names (is aluminium better or worse than bronze? What sits between silver & gold?). Furthermore, if a service is named according to the functionality it provides, then the underlying technology can be swapped out without necessarily having to re-name the service: for example, if a file sharing service moves from fibre-channel disk to SATA disk plus FlashCache, the customer's view of the functionality remains the same, even though the underlying technologies moved from what might once have been called “gold” storage to what might once have been called “silver”.

So consider making bland old “gold”, “silver” and “bronze” and thing of the past, and well-defined, descriptive services that incorporate infrastructure capabilities the way of the future. It will certainly make life less confusing for your projects & customers.

Thursday, 10 January 2013

Stop Thinking About Storage as "Tiers" & Start Thinking About Workloads


In the olden days (which for purposes of this article means "the 1990's"), Gartner introduced - or at least popularised - the concept of "Storage Tiers". The idea was that with new & differentiated storage technologies becoming available, some decision had to be made as to which kind of storage you'd use for a particular kind of data. At the top tier ("Tier 1"), you stored mission-critical, frequently-accessed, latency-sensitive data like OLTP databases. In the middle (call it "Tier 2"), you stored less latency-sensitive data that was still business-critical and needed backup, DR and/or replication such as email systems, and at the bottom (call it "Tier 3"), you stored infrequently accessed or archive data.

Each of these tiers had assumed physical characteristics: Tier 1 was fast, high-performance (and most expensive) disk, usually connected to a fibre channel SAN and featuring high-availability, synchronous replication and so on; Tier 2 was lower-speed disk, with some degree of high availability; Tier 3 was the biggest, cheapest, highest-density disks, probably connected via file sharing networks, and less likely to have high availability features included. These physical characteristics in turn led to an association with specific disk technologies: Tier 1 became small-capacity 15,000rpm fibre channel drives, Tier 2 became large capacity fibre channel drives (maybe operating at 10,000rpm), and Tier 3 became large-capacity SATA drives.

This was fine, for a while: storage arrays frequently provided no additional performance considerations beyond spindle type, spindle size and spindle count; and application managers became used to the the idea that they really only had 3 flavours to choose from, and that Tier 3 was “the slowest”, and Tier 1 was “the best”, so they chose the latter.

Then someone invented solid state drives.

With a tiering system so firmly tied to particular drive technology and “Tier 1” as “the best” (meaning “the fastest”), the introduction a new, faster storage device type created a slight problem: what's better than 1? Fortunately computer scientists all know that you're actually supposed to start counting at zero, so the answer was clear: the new “best” was “Tier 0”, and the top-level descriptions were nuanced to place high-speed transactional data on this new tier.

Problem solved. Until someone invents something faster (like phase-change state drives, or something). At that point we'd have to call the new technology “Tier -1”, which finally clearly shows how ridiculous it is to tie a drive technology to an expected workload.

That's the point of this article – we should be thinking in terms of “workloads”, rather than “tiering”, since tiering is so closely tied to disk technologies, and since the physical drive characteristics are no longer the sole feature to consider. In a NetApp environment there are several features to take into account when designing the solution for a given workload: deduplication & compression, FlashCache, FlashPools, FlashAccel, and so on.

Once we understand what a workload is going to be, we can design a storage system to provide the best combination of features to handle that workload – which may mean that even high-performance workloads are deployed on lower-speed disks. A classic example of this is a typical Virtual Desktop Infrastructure (VDI) workload: many (hundreds or even thousands) of copies of essentially the same operating system and application binary data, with latency-sensitive access. The many copies of data can be deduplicated down to a few or even one instance of actual physically stored data. The first time this data is accessed by a VDI client, it is placed in the controller FlashCache. Subsequent requests for the data from any client are then served directly from the cache. What this means is that the actual disk performance is almost irrelevant, so lower-speed (and lower-cost) drives can be used, and fewer of them thanks to the deduplication effect. The solution becomes cheaper, more efficient, and more performant all at the same time.

This is just one example, and there are plenty more. The main point is that this combination of technologies (slow disk, deduplication, FlashCache) is suitable for that workload, and gives better performance than a traditional “Tier 1” storage infrastructure. It means that it is no longer appropriate to simply use storage tiering to decide the best infrastructure for a given workload. What solution designers need to do now is understand the characteristics of a workload, and then combine the available storage features to most effectively support it.

So from now on, think about workloads, not tiers. This is especially true when trying to develop Infrastructure-as-a-Service offerings. But more about that later.