Showing posts with label Cloud. Show all posts
Showing posts with label Cloud. Show all posts

Saturday, March 07, 2015

Hi Jim ...

A response to Jim's Cloud Post.

---

Hi Jim,

Not quite how I remember the conversation. Factors involved in adoption - efficiency in provision, increase in demand (price elasticity effects, long tail of unmet demand, evolution of higher order systems), rate of innovation of new services (ecosystem effects), ability to take advantage of new sources of wealth (development speed, reduced cost of failure), inertia (16 different forms) & competitor actions. When talking about the shift of infrastructure from product to utility then all these factors come into play. 

Efficiency of Amazon's provision. Do remember that since IT is price elastic and Amazon has a likely constraint e.g. acquiring land and building data centres then Amazon will certainly have to manage its pricing carefully i.e. if it dropped pricing too quickly then demand could exceed supply. So, you need to consider future pricing as well. In all likelihood AWS EC2 is operating at 60%+ margin but this will reduce over time.

Increase in demand. One of most amusing cloud 'tales' is the one that it'll save money. Infrastructure is a million times cheaper today than 30 years ago but has my budget dropped a million fold in that time? No. We don't tend to save money, we tend towards doing more stuff. This is Jevons Paradox. What we need to be mindful of is that our competitors will do more stuff. Which is why you need to be careful about future pricing. If your IT budget is 2% of total budget and your competitor has a 10x advantage then you might shrug it off as a small part of the budget. But, your competitor is likely to end up doing more stuff and suddenly (just to keep up) you'll find you're spending a lot more than 2%.

Rate of Innovation of new service. There are numerous ecosystem games to play in a utility world (such as ILC) which enables a provider to simultaneously be innovative, efficient and customer focused. This seems to be happening with Amazon as all three of those metrics appear to be accelerating. This provides direct benefit for the users of that environment in terms of new service release.

Ability to take advantage of new sources of wealth. Key here is speed and reducing the cost of failure both of which a utility provider offering volume operations of good enough but standard commodity components provides.

Inertia. We all have inertia to change (loss of previous investment, changes in governance / practice, loss of social capital, loss of political capital etc - 16 different forms in total). There will be a counter to any change

HOWEVER ...

The competitive pressure to adapt are often not linear but have exponential effects. If an adaptation gives competitors greater efficiency, increasing access to new services, increases their ability to take advantage of new sources of wealth then as more of your competitors adapt then the  pressure on you mounts. This creates the Red Queen Effect (prof van. valen). 

As a result these forms of change are not linear but exponential. It can take 10 years for such a technology change to reach 3% of the market and a further 5 years to hit 50%. Because of inertia to change and due to its non linear nature many companies (especially competing vendors) get caught out. However, in all such markets there are usually small niches that remain. 

There is also no reason why commoditisation has to lead to centralisation. Many of the forces can be countered. Unfortunately due to the incredibly sucky play of often past executives within competitors then in this case centralisation (to AWS, MSFT and Google a distant third) seems very likely. Some of those past executives were warned in 2008 about how to fragment the market by creating a price war with AWS clones forcing demand beyond Amazon's ability to supply due to the data centre constraint. It's shocking that they were so blinded that they've got large companies into this state.

So, 

1) Will infrastructure centralise to those players of AWS, MSFT and GooG? Yes plus clones of those environment. Competitors have shown pretty poor strategic play in the past and this is now the most likely outcome.

2) Will everything go to public infrastructure clouds? No. There will be niches. There is also inertia to the change but the pressure will mount (Red Queen) as competitors adopt cloud. The change usually catches people out due to its exponential nature.

3) Is it just price? No. Multiple factors involved. Price is one of those factors.

4) Why is Walmart building a modest sized private cloud? Probably because it's concerned over Amazon's encroachment into its own retail industry. In all likelihood they will end up adopting Azure or GooG over time.

Wednesday, February 25, 2015

What's wrong with my private cloud ...

These days, I tend not to get too involved in cloud (having retired from the industry in 2010 after six years in the field) and my focus is on improving situational awareness and competition  (in my view a much bigger problem) through techniques such as mapping.  I do occasionally stick my oar into the cloud due to some very elementary mistakes that appear. One of those has raised its head again, namely the cost advantages / disadvantages of private cloud.

First, private cloud was always a transitional play which would ultimately head to niche unless a functioning competitive market forms enabling a swing from centralised to decentralised. This functioning market doesn't exist at the infrastructure layer and so the trend to niche for private cloud is likely to accelerate. There's a whole bunch of issues related to the impact of ecosystems (in terms of innovation / customer focus and efficiency rates), performance and agility, effective DevOps, financial instruments (spot market, reserved), capacity planning / use, focus on service operation and potential for sprawl which can work out very negatively for private cloud but these are beyond the scope of this post. I simply want to focus on cost and to make my life easier - the cost of compute.

Here's the problem. Two competitors going head to head in a field, both with over £15Bn in annual revenue and both spending over £150 M p.a. on infrastructure (i.e. compute, hosting etc). Now, this IT component represents a good but not huge chunk of their annual revenue (1%).

One of the competitors was building a private cloud. It was quite proud of the fact that it reckoned it had achieved almost parity with AWS EC2 (taking in this case a comparison to a m3.medium) of around $800 per year. What made them happy is that the public Amazon cost is around $600 per year and so they weren't far off but they gained all the "advantages" of being a private cloud. This is actually a disaster and was caused by three basic mistakes (ignoring all the other factors listed above which make private cloud unattractive).

Mistake 1 - Externalities.

The first question I asked (as I always do) is what percentage of the cost was power? The response was the usual one that fills me with dread - "That comes from another budget". Ok, when building a private cloud you need to take into consideration all costs from power, building, people, cost of money etc etc. On average the hardware & software component tends to be around 20-25% of the cost. Power tends to be the lion share. So, they weren't operating at anywhere close to $800 per equivalent per year but instead closer to $3,000 per year. This on the face of it was a 5x differential but then there's - future pricing.

Mistake 2 - Future pricing

Many years ago I calculated how efficiently you could run a large scale cloud and from this guesstimated that AWS EC2 was running at over 80% margin. But how could this be? Isn't Amazon the great low margin, high volume business? The problem is constraint. 

AWS EC2 is growing rapidly and compute happens to be elastic i.e. the more we reduce the price, the more we consume. However, there's a constraint in that it takes time, money and resource to bring large scale data centres online. With such constraints it's relatively easy to reduce the price so much that demand exceeds your ability to supply which is the last thing you want. Hence you have to manage a gentle decline in price. In the case of AWS I guesstimated that they were focused on doubling total capacity each year.  Hence, they'd have to manage the decline in price to ensure demand kept within the bounds of their supply factoring in the natural reduction of underlying costs. There are numerous techniques that can also help i.e. increasing size of default instance etc but we won't get into that.

Though their price is currently $600 per year, I take the view that their costs are likely to be sub $100 which means that a lot of future price cuts are on the way.  The recent 'price wars' from Google seem more about Google trying to find where the price points / constraints of Amazon are rather than a fully fledged race to the bottom. All the competitors have to keep a watchful eye on demand and supply.

Let us assume however, that AWS is less efficient than I think and the best price they could achieve is $150. This suddenly creates a future 20x differential between the real cost of the private environment. However, it's no big shakes because even if the competitor was using all public cloud (i.e. data centre zero) which is unlikely then it simply means they're spending $7.5M compared to our $150M and whilst they might be $140M of saving this is peanuts to our revenue ($15Bn+) and the business that's at stake. It's not worth the risk.

This couldn't be more wrong.

Mistake 3 - Future Demand

Cloud computing simply represents the evolution of a world of products to a world of commodity and utility services. This process of "industrialisation" has repeated many times before in our past and has numerous known effects from co-evolution of practice (hence DevOps), rapid increases in higher order systems, new forms of data (hence our focus on big data), increases in efficiency, a punctuated equilibrium (exponential change), failure of companies stuck behind inertia barriers etc etc.

There's a couple of things worth noting. There exists of long tail of unmet business demand for IT related projects. Compute resources are price elastic. The provision of commodity (+utility) forms of an activity enable rapid development of often novel higher order systems (utility compute allows a growth in analytics etc) which in turn evolve and over time become industrialised themselves (assuming they're successful). All of this increases demand for the underlying component.

Here's the rub. In 2010 I could buy a million times more compute resource than in the 1980s but does that mean my IT budget for compute has reduced a million fold in size during that time? No. What happened is I did more stuff.

In the same way, Cloud is extremely unlikely to reduce IT expenditure on compute (bar a short blip in certain cases) because we will end up doing more stuff. Why? Because we have competitors and as soon as they start providing new capabilities then we have to co-opt / adapt etc.

So, let us take our 20x differential (which in all likelihood is much higher) and assume we're building our private cloud for 5yrs+ giving time for price differentials to become clear. Our competitor isn't going to reduce their IT infrastructure spending, they are likely to continue spending $150M p.a. but do vastly more stuff. However, in order to maintain parity with this given the differential then we're going to need to be spending closer to $3Bn p.a. In reality, we won't do this but we spend vastly more than we need and we will still just lose ground to competitors - they'll have better capabilities than we do.

You must keep in mind that some vendors are going to lose out rather drastically to this shift towards utility services. However, they can limit the damage if you put yourself in a position where you need to buy vastly more equipment (due to the price differential) just to keep up with your competitors. This is not being done for your benefit. When you're losing your lunch, sometimes you can recover by feasting on a few individuals. If that means playing to the inertia of would be meal times (FUD, security, loss of jobs etc) then all is fair in war and business. You must ask yourself whether this is the right course of action or am I being lined up as someone's meal ticket?

Now, I understand that there are big issues with heading towards data centre zero and migrating to public cloud often related to the legacy environments. This is a major source of inertia particularly as architectural practices have to change to cope with volume operations of good enough components (cloud) compared to past product based practices (e.g. the switch from scale up to scale out or N+1 to design for failure etc). We've had plenty of warning on this and there's all sorts of sensible ways for managing this from the "sweat and dump" of legacy to (in very niche cases) limited use of private. 

But our exit cost from this legacy will only grow over time as we add more data. We're likely to see a crunch on cloud skills as demand rockets and this isn't going to get easier. There is no magic "enterprise cloud" which can enable future parity with all the benefits of commodity and volume operations but provided with non-commodity products customised to us in order to make our legacy life easier. The economics just don't stack up. 

By now, you should be well on your way to data centre zero (several companies are likely to reach this in 2015). You should have been "sweating and dumping" those legacy (or what I prefer to call toxic) IT assets over the last four+ years. You should have very limited private capacity for those niche cases which you can't migrate unless you can achieve future pricing parity with EC2 (be honest here, are you including all the costs? What is your cost comparison really?). You should have been talking with regulators to solve any edge cases (where they actually exist). You should be well on the way to public adoption.

If you're not then just hope your competitors return the favour. Be warned, don't just believe what they say but try and investigate. I know one CIO who has spent many years telling conferences why their industry couldn't use public cloud whilst all the time moving their company to public cloud. 

Some companies are going to be sorely bitten by these three mistakes of externalities, future pricing and future demand. Make sure it isn't you.

-- Additional note

Be very wary of the hybrid cloud promise and if you're going down that route make sure it translates to "a lot public with a tiny bit of private". This is a trade off but you don't want to get on the wrong side and over invest in private. There's an awful lot of vendors out there encouraging you to do what is actually not in your best interest by playing to inertia (cost of acquiring new practices, cost of change, existing political capital etc). There's some very questionable "consultant" CIOs / advisors giving advice based upon not a great deal and dubious beliefs.

Try and find someone who knows what they're talking about and has experience of doing this at reasonable scale e.g. Bob Harris, former CTO for Channel 4 or Adrian Cockroft, former Director Engineering Netflix. These people exist, go get some help if you need it and be very cautious about the whole private cloud space.

-- Additional note

Carlo Daffara has noted that in some cases, the TCO of private cloud (for very large, well run environments) can achieve 1/5th of AMZN public pricing. Now, there are numerous ecosystem effects around Amazon along with all sorts of issues regarding variability, security of power, financial instruments (reserved instances etc) and many of those listed above which create a significant disadvantage to private cloud. But in terms of pure pricing (only part of the puzzle) then a 1/5th AMZN public pricing is just on the borderline of making sense for the immediate future. Anything above this and then you're into dangerously risky ground. 

Wednesday, December 07, 2011

Future costs and Cloud

There are many subjects which I find tiresome but two which are starting to irritate me are the notion of Enterprise Cloud and Financial ERP. I'll deal with Enterprise Cloud in this post.

The shift from products to utility services inevitably incurs various forms of risks. These include disruption risks such as loss previous skillsets and political capital to transitional risks such as changes to governance and transparency of suppliers to outsourcing risks such as pricing competition and loss of strategic control.

A common, past method of dealing with transitional risks is the use of a hybrid model combining both public and private supplies. However, this is a transitional approach and should be undertaken with a view of moving to a future hybrid model of multiple public providers (i.e. a competitive market).

A transitional approach requires you to build in a way which is likely to be compatible with a future public market. For infrastructure this mean use of commodity components and in most scenarios an EC2 / S3 / EBS like interface. Hence my general support for open efforts like OpenStack (and to a lesser extent Eucalyptus).

Unfortunately, many applications are designed with the best practice for a product world i.e. scaling is about bigger machines, resilience is about N+1 and in general the focus is on reliable hardware. Best practice for a utility world involves resilience, scaling and failure modes built around software i.e. design for failure, distributed systems and chaos engines such as Netflix's chaos monkey approach. There is an inevitable cost of architectural transition from one set of best practices to another.

Obviously many companies don't like this architectural cost of change and hence want to minimise it which has given rise to the concept of the Enterprise cloud i.e. it's like cloud but without the commodity bit.

It should be noted that cloud is simply a result of a standard process of evolution that inevitably leads to operational efficiency through provision of a commodity. A consequence of this is it also enables higher rates of innovation for new business activities (such as big data) through the combined effects of componentisation and creative destruction. The upshot of this, is that you've never had a choice with cloud - it's just a question of when and the longer you leave it then the more you put yourself at a competitive disadvantage to others.

Whilst a commodity based private cloud (which can and should achieve much lower costs than public provision today) is a viable option in the short to mid term (depending upon scale), unfortunately Enterprise clouds don't move you along that architectural transition and here there's a real gotch'a. The problem is simply known as Jevons' paradox.

As competitors gain the benefits of more efficient commodity provision and higher rates of creation, this is unlikely to result in a reduction in IT budgets but instead more IT activities undertaken with everyone trying to keep up with each other. We've seen this for the last thirty years i.e. as IT has become more efficient, IT budgets haven't fallen but we've just ended up doing more stuff.

You therefore have to factor in that five to six years from now your architectural transition costs may well have spiralled by an order of magnitude due simply to the increased size of the estate. This can obviously be counter balanced with a simplification strategy, assuming your estate is already bloated but when considering Enterprise cloud, you must include this increase of architectural transition costs along with less efficient provision during that time due to a non commodity approach. Even private clouds are going to look dubious in this timeframe.

In most cases, Enterprise cloud will be a pretty unattractive option with ongoing and increasing costs.  It can however still be useful as part of a sweat and dump strategy for legacy environments i.e. you push the capital costs for legacy onto a provider with a view of dumping that part of the estate in the near term.

Why do I find this subject irksome? I simply hate repeating old ground and this has been covered many times before over many years. This is the last time.

-- 26th Nov 2013

Thursday, October 06, 2011

Larry offers Hotel California ...

According to CW, Larry recently raged over Salesforce describing it as a roach motel and pleading the case for interoperability. He's just given a gift horse to some fairly smart operators and this time Larry's forgotten to fill it with any of his own soldiers.

First, some background which everyone knows already, so I'll keep it short :-

  • Activities evolve and our industry has been shifting from a product to a utility service world. This has been clear for the last 6+ years, Salesforce knows this and they've been positioning themselves in that future space.

  • Past success always acts as an inhibitor to future survival, it creates an inertia barrier to change. This is why Amazon and not some hosting company encumbered by an existing business model made the break into IaaS. This has been crystal clear for 4+ years. Salesforce knows this, it's why Oracle has been slow to react to the change.

  • In this future world, competitive markets will become key to solving those outsourcing risks such as pricing competition, second sourcing options and loss of strategic control. Such markets will require multiple providers, access to code and data (i.e. standard data formats and APIs) and semantic interoperability. The latter point is only solvable with complex systems through running code and unless the market intends to be a captured markets (i.e. dependent upon one vendor) then that code will have to be open source. This has been clear for the last 5+ years. Everyone knows this just a lot of people refuse to believe it usually because of inertia barriers which have become institutionalised.

  • Critical in this new world is the development of ecosystems as these enable a company to solve the innovation paradox and simultaneously appear more innovative and highly efficient. This has been blindingly obvious for 3+ years. More details on common models such as ILC can be found here. Salesforce knows this, they've been playing an acquisition game around their own ecosystem and sending market signals because of this.
  • With a large enough ecosystem, you can create network effects through aggregated data e.g. market reports. This can be used as a soft form of lock-in i.e. even if you open source an entire system, your service still maintains an advantage simply because of the number of people using it. In other words, you can be entirely open but in effect create lock-in (i.e. gravity) for your service because of the benefits that being within that ecosystem brings. This has been painfully obvious for the last 3+ years.
  • Salesforce has also been playing a tower and moat ploy, building a tower of core revenue surrounded by a moat of high barriers to entry and devoid of differential value. Attacking Salesforce is a tough call for anyone, hence I suspect Larry's aim to make interoperability his calling card.
Salesforce has the ecosystem to play an aggregated data game i.e. free market reports for an industry based upon aggregated data or free comparison KPIs to your sales team effectiveness etc. Given the smart plays Salesforce has been making, you can bet your bottom dollar they've got lots of this in the pipeline.

Salesforce could also use open source as a tactical weapon in this space. They could open source the entire system and say "come and compete", "run it yourself" with full knowledge that those who build it for themselves and take the private road will eventually switch to public, whilst those setting up as public providers will lack the ecosystem and hence any aggregated data benefits. Salesforce is also smart enough to know that this game could be played against them, so they'll have to go down that route at some point. Hence you can pretty much bet your bottom dollar they've been working on this.

Larry has walked into a huge trap. He's just called out interoperability as the key differentiator for his service but as we all know the real issue is portability which requires semantic interoperability and running code. All Salesforce has to do is start launching more aggregated data services and open source the entire system under a banners of "Freedom in the cloud", "Run it yourself for Free" and Larry is left standing with the high cost proprietary service with no real portability (except between one licensed version of Oracle and another).

It's difficult to see how Oracle's strategists could have been more tweedledum or tweedledee as currently they are primed to become the industry's example of Hotel California (you can go anywhere you like as long as you're paying fees to Oracle?).

Now, open sourcing won't be easy for SFDC because they have an existing service, security professionals will be concerned over exposing security weaknesses, lawyers will have their usual collywobbles over IP and financial controllers will gasp at writing down a technology asset.

However Benioff like Maritz (you don't think CloudFoundry doesn't have a grand strategic purpose do you?) is generally a shrewd player. It all boils down to a question of timing and willingness to play the end game but we could be expecting checkmate to Salesforce in the near future.

Bad move Larry ... really bad. Oracle will be lucky to make it out of 2020 with this standard of play.

Friday, September 02, 2011

The battle that wasn't ...

Chris Boos wrote an interesting post about a debate that @samj and I were having on twitter regarding APIs in the cloud space. I thought I'd leave my comment here as a general view on the subject.

A couple of things to point out. Twitter is not the best tool in the world to determine the exact context of a discussion because those listening aren't generally privy to the history of the discussion. Hence in this case, it may not be clear that Sam Johnston and I are in absolute agreement on the importance of open source and efforts like OpenStack in this world.

Any difference between us is on the necessity of reverse engineering APIs and co-opting as the main short term tactical play. The long term we're both totally in agreement on - open standards, open formats and open source are critical.

Our difference in views on short term tactical plays hardly constitutes a battle but is merely debate. As for being a "giant", whilst that is very flattering it doesn't coincide with my view of the world. Nevertheless, it was an excellent post by Chris and much appreciated.

Comment --

Just to clarify my view - as it currently stands any company can reverse engineer an API for reasons of interoperability. Hence when trying to make a market of providers in the IaaS space with semantic interoperability between providers, I strongly support adoption where there is clearly a dominant API.

It should be noted that such a market can have multiple open source and proprietary implementations around the API. However, running code through an open source effort is necessary to form a market place without a single (or consortium of) vendor(s) being able to force a tax on that market. In other words, providers need to have an operational means of implementing the service and compete in the market without a necessity to purchase software licenses (a tax on competition). They may choose to buy software to do so but a free market is one unencumbered by such forced taxation.

This is why I do no support MSFT Azure's effort, despite the provision of open standards because there exist no open source implementation.

This is why I did not support Google's AppEngine, despite the provision of an SDK as there existed no fully operational open source means of implementing the service.

This is why I strongly support open source efforts which reverse engineer the dominant API for reasons of interoperability e.g. open stack, eucalyptus etc.

It is also why I strongly support open source efforts which attempt to create the dominant standard in a fledgling market, such as CloudFoundry in the PaaS arena.

Once the marketplace of alternative providers is large enough and it has the dominant ecosystem then the open source effort in effect becomes the defacto standard for implementation and the API in that market. If necessary, due to abuse of position by the original provider, then the API can be differentiated away from the original provider including providing an entirely new API where applicable.

I don't find attempts to differentiate on API in a utility world where one API is clearly dominant meaningful. Of course if an open source effort (such as openstack) creates a large enough ecosystem then it is in effect the dominant and can do as it pleases.

I find re-inventing the wheel by creating an API by committee and attempting to get the market to adopt as a wasted effort when a market has in principle chosen.

I do find the way to standardise is through creating the largest ecosystem and in such cases both reverse engineering the dominant API for reasons of interoperability combined with provision of open source running code is necessary.

Co-opt rather than compete is the order of the day in this world.

Thursday, August 18, 2011

Hosting Con Keynote

I was very fortunate to be asked to give the opening keynote at Hosting Con 2011 covering commoditisation, business evolution, leadership and what the various tactical plays in the cloud computing space mean to hosting companies. The audience was fantastic, I had a great time and despite using excessive numbers of slides, no-one was hurt in the process.

Continuing on the theme from my OSCON tutorial, I've uploaded a summary set of slides which are highly condensed but give a taster to what we covered.

Alas, there's no video and as per usual I'm six years into writing my book and around 30% of the way there. The subject matter keeps on giving me more areas of interest to explore, so don't hold your breath for me to finish any time soon.

Tuesday, August 02, 2011

OSCON Tutorial

I gave a three hour tutorial at OSCON on innovation, commoditisation, business evolution, organisation, leadership and various tactical plays in the cloud computing space. The talk was a blast, I really enjoyed it and judging by the feedback it hit some home runs with many of the audience.

However, the presentation is 1,041 slides long and so - I'm not uploading that or creating a video. Instead I've made a summary presentation which covers the main points.

Be warned, it's highly condensed.

Tuesday, April 12, 2011

Open source as a tactical weapon, VMware's latest move.

All business activities evolve through a common lifecycle and we're currently witnessing a shift of many IT related activities from a product to a utility service world. This is commonly referred to as "the cloud". This transition brings benefits, risks, different methods of operating but also impacts the tactical plays in the great skirmish between companies. There are two models of tactical play which are particularly noteworthy - ILC and Tower & Moat.

In my previous LEF post, I discussed the innovate, leverage and commoditise (ILC) model that seems to be naturally appearing in companies such as Salesforce. To summarize, it is a technique by which a company uses a surrounding ecosystem to not only reduce the cost of innovation but to encourage innovation and rapidly identify success. By acquiring and providing such innovations as common services, a virtuous circle can be created.

A second model is the tower and moat. The principle here is to defend a revenue stream (the tower) by creating a moat devoid of differential value with high barriers to entry around it. By way of example, Salesforce created its own tower around provision of CRM as a utility service in a world where CRM was generally provided through customisable products or rental services. As barriers to entry into this new field were eroded (i.e. Amazon enabling widespread access to utility infrastructure) then new barriers were created through the acquisition of platform technology.

Whilst product based competitors attempt to differentiate themselves with activities such as social CRM, Salesforce acquired such activities with the view of providing common services. The net effect is this eliminates the differential value of social CRM and helps establish a moat. Salesforce has been extensively using its ecosystem (an ILC model) to identify and acquire a wide range of potential differentials and further strengthen its moat.

When competitors finally move to a cloud model then they will find the space inhabited by a large player with a large ecosystem and few opportunities to differentiate - a reasonably fatal combination. Both ILC and the Tower & Moat model are powerful tools which can also be used to counter competitors. They can be used together, or individually or combined with other tactical plays such as open source.

Take the case of Apple vs Android : whilst the iPhone is not one activity but a device describing many activities, Google has effectively created an ecosystem around Android which provides a means of identifying and accelerating innovation in this field whilst reducing costs. By providing the system as open source and creating a hardware ecosystem, then Android has effectively removed much of the differential value that Apple might have sort. Apple would appear to have been pushed into a high risk, stand alone innovation game against a broad ecosystem.

Take the case of cloud infrastructure : we've already seen Rackspace & NASA move to create open source software - the OpenStack project - to provide infrastructure as a service. Their vision is to create a competitive marketplace of computer utilities around OpenStack. Such a world plays to Rackspace's strength of service delivery as a utility provider but also fits with NASA's goals of increasing efficiency of infrastructure. The ecosystem around openstack should encourage rapid innovation and if successful will create the standard that a competitive marketplace depends upon. It will also drive out differential value in this space making it tough for new competitors or those with a proprietary offering. I say 'should' and 'if' because I have real concerns over the differentiation from Amazon idea.

Take the case of large scale infrastructure: into which Facebook has announced the OpenCompute project and in effect open sourced how to create large scale data centres. This should over time help eliminate differential value that such knowledge created and whilst beneficial to the future computer utility world it will also help to undermine those for whom such skills have acted as a barrier to entry into their industry - namely massive scale search engines and data processors.

Take the case of healthcare: which has seen the VA (Veterans' Association) create an open source electronic health record system from VistA. It seems clear that the VA are focused on encouraging innovation through ecosystem effects and creating a marketplace of competitive providers. Visions of a worldwide standard are not beyond the realm of reason.

Take the case of platform as service into which VMware has announced an integrated set of open source platform components known as CloudFoundry. If successful and there's every reason to believe it will be then VMware will succeed in creating a huge moat devoid of differential value in the platform space and a vast ecosystem driving this. Any would be competitors will face an uphill struggle to compete against VMware's effort. Those planning proprietary platform offerings should take note of this move.

But wait ... where's the tower?

The beauty of creating a competitive marketplace of utility service providers is that it opens up a huge range of opportunities from service provider, support, assurance, brokerage, exchange, marketplace and a dozen more. Being at the heart of this, which is where VMware will be, means they are well positioned to take advantage. It's a bold move, perfectly timed and well executed.

Of course, CloudFoundry has already been made to run on Amazon EC2 which means CloudFoundry on OpenStack built on an environment designed around OpenCompute can't be far behind.

The world of IT is changing and many IT activities have become suitable for provision through utility services. With this change comes tactical plays designed to take advantage of this shift. At the OSCON conference in July 2007, I stated that in this future utility world open source was the only way of effectively competing. Time will tell but the increasing drive towards open source and its use by major companies as a tactical weapon seems to be pointing that way.

Smart move by VMware, it'll certainly shake up the industry.

--- Update 5th May 2014

Most is proceeding as expected. Bizarrely SAP / Oracle just seem to be waking upto the threats ... a bit too late. Unfortunately also OpenStack continued its differentiation play and the market never formed. Cloud Foundry however is storming ahead. Apple is starting to look weak vesus Android whilst OpenCompute gathers momentum.

Monday, April 11, 2011

A question of standards.

I'm all in favour of standards for the "cloud" world because such standards should enable and accelerate innovation in IT through componentisation effects as well as encourage the formation of competitive markets of compute utilities

However, that said, I'm against standards committees and the concept that open standards (as in APIs & data formats) are enough to create the portability required for competitive markets.

In the former case, the market will decide the standard and the job of any standards body should be to rubber stamp what is an existing practice. Unfortunately, standards committees are often used as vehicles to promote specific vendor interests and in many cases their efforts are counter productive. For example, whilst IPX/SPX was the committee approved standard, it was TCP/IP which won the marketplace battle. The only impact that approving IPX/SPX as a standard had was to temporarily slow adoption of TCP/IP in some quarters and hence inhibit innovation.

Standards should have a positive effect but defacto has to precede dejeure. There are some exceptions to this but they are exceptions.

In the case of open standards, a competitive marketplace requires multiple providers, access to code and data (ideally with syntactic interoperability) and semantic interoperability of services. Whilst open standards provide part of the solution, it is critical for reasons of semantic interoperability that a common reference model (i.e. running code) is provided.

Since, we're talking about a world of service (and not feature) differentiation for activities which are fundamentally commodity by nature and hence suitable for utility service provision, then the obvious solution is an open source reference model as the standard. Potential examples of such would be the OpenStack effort.

Such open source reference models would ideally exploit the dual nature of GPLv3 which is both restrictive in the product world but simultaneously permissive in the service world to create a functioning marketplace. Unfortunately, many vendors promote open standards (as in APIs etc) as the solution to these problems of portability and hence describe their systems as open when they're quite clearly not.

So, in general :-

  • Are standards good for the cloud?
    Absolutely, it'll encourage innovation through componentisation effects.

  • Are standard committees good for the cloud?
    Generally no. At best they should rubber stamp market chosen approaches however in reality they're more likely to get in the way or slow progress by promoting vested interests.

  • Is portability between providers important for the Cloud?
    Absolutely, it's the route to formation of competitive marketplaces and reducing outsourcing risks.

  • Will open standards provide the portability needed?
    Only in the most trivial cases but not for the vast majority of activities. Open standards are necessary but they are not sufficient to provide the portability required. The idea that open standards alone will achieve this will inhibit the formation of competitive markets.

  • Is open source essential for the cloud?
    Absolutely, the formation of competitive markets without loss of strategic control to a specific vendor depends upon the provision of open source reference models as the standard.

Monday, February 28, 2011

Private vs Enterprise Clouds

There is a debate raging at the moment between different types of clouds hence I thought I'd stick my oar in. First, let's clear up some simple concepts:-

What is cloud?
Cloud computing is simply a term used to describe the evolution of a multitude of activities in IT from a product to a utility service world. These activities occur across the computing stack.

These activities are evolving because they've become widespread and well defined enough to be suitable for utility provision, the technology to achieve this utility provision exists, the concept of utility provision is widespread and there has been a sea-change in business attitude i.e. a willingness to accept these activities as being a cost of doing business and commodity-like.

There is nothing special about this evolution, it's bog standard and many activities in many industries have undergone this change. During this change, the activities are moving from one model to another which creates a specific set of risks (both real and perceived) around trust in the new providers, transparency, governance issues and concerns over security of supply. There are other risks (including outsourcing and disruption risks) but these are irrelevant for the specific question of private vs enterprise clouds. Nevertheless, all of these "risks" help increase inertia to the change. For clarity's sake, I've summarised this evolution in figure 1.

Figure 1 - Lifecycle and Risk.
(click on image for larger size)


Why private clouds?

Private clouds (where the service is dedicated i.e. "private" to a specific consumer) are generally used as part of a hybrid strategy which combines multiple public sources with a private source. The purpose of a hybrid strategy is simply to mitigate transitional risks (concerns over governance such as data governance, trust etc). This is a normal supply chain management tactic and occurs frequently with this type of evolution. Even within the electricity industry you can find a plethora of hybrid examples in the early days of formation of national grids.

Fundamentally, a hybrid strategy mitigates risk but incurs the costs of both lesser economies of scale and additional resource focus when compared to a pure public play. It's a simple trade-off between benefits and risks which can often be justified for a time.


Why Enterprise Clouds?

Enterprise clouds need a bit more explaining and to understand why they exist we first have to start with architecture. To keep things really simple, I'm going to focus on infrastructure.

As per above, the use of computing infrastructure has undergone a typical evolution from innovation to commodity forms and more recently the appearance of utility services. In the earlier stages of evolution, the solution to architectural problems was bound to the physical machine. Scaling was solved through buying "a bigger machine" i.e. Scale-Up whereas resilience involved hot-swap components, multiple PSUs and other redundant physical elements (the N+1 model). This was essential because of the long lead time for replacement of any physical machine i.e. there existed a high MTTR (mean time to recovery) for a physical server. Applications therefore developed on the assumption of ever bigger and ever more resilient machines. These architectures spread (specific sets of knowledge behave just like activities) and became best practice.

As computing infrastructure became commodity-like, novel system architectures such as Scale-Out (aka as horizontal scaling) developed. These solved scaling by distributing the application over many more smaller and standardized machines. The new architectural solution was therefore "buy more machines" and not "buy a bigger machine". This scale-out architecture rapidly spread as infrastructure became more generally accepted as a commodity.

As we entered the utility phase, infrastructure has in effect become code – that is, we can create and remove virtual infrastructure through API calls. The MTTR of a virtual machine provided through a utility service is inherently lower than its physical counterpart and novel architectures called "design for failure" have emerged that exploit this. The technique involves monitoring a distributed system and then simply adding new virtual machines when needed.

In the cloud world application scaling and resilience are solved with software whereas in legacy it was often solved by physical means. This is a huge difference and I've summarized these concepts in figure 2.

Figure 2 - Evolution of Architectures
(click on image for larger size)


By necessity, public cloud infrastructure is based upon volume operations – it is a utility business after all. Virtual compute resources come in a range of standard sizes for that provider and are based upon low cost commodity hardware. These virtual resources are typically less resilient than their top of the line physical enterprise class counterpart but naturally they are vastly cheaper.

It should be noted that a ‘design for failure’ approach can take advantage of these low cost components to create a far higher level of resilience at any given price point through software.

By way of example, the general rule of thumb is that each 9 roughly doubles the physical cost. Hence let's take a scenario of a base machine designed to give a 99% up-time with a machine designed to provide 99.9% uptime costing twice the amount etc.

In the above scenario, using four base machines in a distributed architecture provides a theoretical up-time of 99.999999% though naturally it will suffer periods of degraded performance due to single, dual or triple machine failure. In a utility computing environment, this impact is negligible due to the low MTTR of creating new virtual machines and in such as case you have a low cost environment (4x base unit), high level of resilience for total failure (1 in 100 million) and fast recovery for periods of degraded performance due to the low MTTR.

Now compare this to a single physical machine (ignoring all scaling, network and persistence issues etc). The equivalent physical machine would be 16x more costly (assuming the 2x rule holds) and in the rare situation that complete system failure occurred, the MTTR would be high. In reality, MTTR would be high for all its components unless spares were kept.

Now for reasons of brevity, I'm grossly simplifying the problem and taking a range of liberties with concepts but the principle of using vast numbers of cheap components and providing resilience in software is generally sound. We've seen many examples of this including RAID. These cloud architectures are no different except they extend the concept to the virtual machine itself.

The critical point to understand is that these two extremes of architecture have fundamentally different models based upon the differences in the underlying components i.e. the architectural practice co-evolved with the activity. The cloud model is based upon volume operation provision of commodity good enough components with application architectures exploiting this through scale-out and design for failure. Now in the following figure I've tried to highlight this difference by providing a greyscale comparison between a traditional data centre and a public cloud on a number of criteria.

Figure 3 - Data Centre vs Cloud.
(click on image for larger size)


OK, now we have the basics let's look at the concept of Enterprise Cloud. The principle origin of this idea is that whilst many Enterprises find utility pricing models desirable, there is a problem when shifting legacy environments to the cloud. Now when companies talk of shifting an ERP system to the cloud they often mean replacing own infrastructure with utility provided infrastructure. The problem is that those legacy environments often use architectures based upon principles of scale-up and N+1 and infrastructure provided on a utility basis doesn't confirm to these high levels of individual machine resilience. The problem is simply people are trying to shift applications built with best architectural practice for product based infrastructure to a world where infrastructure is a utility which has its own but different best practice. It's not that legacy is wrong, it's just that in this case legacy means built with best practice for a product world and that practice is no longer relevant.

At which point the company has two options; either re-architect to take advantage of the volume operations through scale-out and design for failure concepts or demand for higher level resilient virtual infrastructure i.e. try and make the new world act like the old world.

This is where Enterprise Cloud comes in. It's like cloud but the virtual infrastructure has higher levels of resilience but at a high cost when compared to using cheap components in a distributed fashion. So why Enterprise cloud? The core reason behind Enterprise Cloud is often to avoid any transitional costs of redesigning architecture i.e. it's principally about reducing disruption risks (including previous investment and political capital) by making the switch to a utility provider relatively painless. However, this ease comes at a hefty operational penalty. The real gotcha' is those transitional costs for redesign increase over time as the system itself becomes more interconnected and used.

In principle, Enterprise Clouds are used to minimise disruptional risks but they will be ultimately subject to a transition to the new architecture because of high operational costs. There are other reasons for an "Enterprise Class" cloud but most of these, such as where data resides, can also be provided by using a "Private" cloud that is built using commodity components until suitable "Public" clouds are available.

The different economic models is what separates out private / public compute utilities from enterprise clouds, I've highlighted this on the greyscale in figure 4. Actually, I prefer to use the original term virtual data centre rather than enterprise cloud because that's ultimately what we're talking about.


Figure 4 - Enterprise vs Private Cloud.
(click on image for larger size)


Do Enterprise Clouds have a future? Sort of, but they'll ultimately & quickly tend towards niche (specific classes of SLAs, network speeds, security requirements etc). Their role is principally in mitigating disruption risks but the transitional costs they seek to avoid can only be done at an increasing operational penalty. It should be noted that there is a specific tactic where an enterprise cloud can have a particularly beneficial role: the "sweating" of an existing legacy system prior to switching to a SaaS provider (i.e. "dumping" the legacy). In most cases, Enterprise clouds will become an expensive folly.

Do Private Clouds have a future? Yes, they have a medium term (i.e. next few years) benefit in mitigating transitional risks (such as issues over data governance) through the use of a commodity based model. However, be careful what you're building and remember the impact of private clouds will diminish as the public markets develop. I say "be careful what you're building" but in practice this means don't build. The vast majority of companies lack the skills and management capability necessary and though you're trying to build a "commodity" based private cloud, the chances are you'll end up with some very "enterprise cloud" like.

Do Public Clouds have a future? Absolutely, this is the long term economic model and will become the dominant mechanisms. Public utilities should also encourage clearing houses, exchanges and brokerages to form. 

Last thing, I said I would focus on infrastructure to keep things simple. The rules change when we move up the computing stack ... but that's another post, another day.


--- 8th February 2016

Note to self.

Five years later and ... oh, my past self would not want to know.

People are still building private clouds and trying to push enterprise class cloud despite the obviousness of the changes. The whole co-evolution of practice finally became firmly established as DevOps but has gone a little bit out of control in terms of making somewhat grand claims of cultural changes. The points of inertia and the overstating of risks continues in corporations but we're finally getting through that point in the punctuated equilibrium that this will all get washed away.

Overall, it has been a torrid journey. Many are still confused on basic concepts such as the change of practice with activities. We're still having this discussion today! I have to caution that legacy isn't flawed as much as built with best practice for a product world. Many are trying to "create" new paradigm shifts in order to re-establish flagging business models by "having another go". In many cases, the strategic play of many former IT giants has been next to hopeless. There's quite a list of well known names circling the spiral of death by cost cutting to restore profitability whilst the underlying revenues are in decline.

It has become plainly clear that the level of blindness to the environment (i.e. poor situational awareness) is incredibly high in most corporations at the executive level and far worse than I could have possibly anticipated. Inertia will always need to be managed but that such companies failed to manage predictable changes with so much warning is ... stunning. 

Saturday, February 26, 2011

Will Cloud Computing help the business align to the market?

I was recently at a private conference of CIOs when in the midst of discussing activities such as CRM to ERP, it became clear that not only did everyone understand these applications, but also that every company had them. Almost all of those companies have expensive customization programmes in place, tailoring these systems to fit their needs. This behaviour is puzzling, but to explain why, we'll need to first look at the concept of business evolution.

All business activities share a common evolutionary pathway, a lifecycle; from the innovation of an activity, to custom built systems implementing that activity, to the provision of products, to eventually commodity provision and the appearance of utility services.

It is unusual to find activities that are commonplace, well understood, well defined and accepted as a cost of doing business being treated as though they were in an earlier stage of evolution with heavy customization. But each CIO told how that was exactly what was happening at their company, and the discussion became surreal when we discovered that many of the customizations being made to systems like CRM were common as well.

Surely, these activities are best served by being provided as a standardized commodity ideally through a market of utility services? Isn't that the whole point of cloud computing? So why aren't we all flocking to consume a host of common activities through the utility services of cloud providers? Well, a diverse range of companies already are.

However, several of the CIOs who had looked into the issue said these services didn't work for their company. On further questioning, it became clear that it wasn't IT but the business that was pushing for the customization in order to fit in with their way of working. Several CIOs had even used the standardized services of cloud providers to challenge this - asking the business to justify the additional orders of magnitude costs for a customized system by demonstrating the differential value that 'their way' made.

Whilst we often talk about business and IT alignment, this shouldn't mean IT just delivering what the business asks for. In circumstances where there is no differential value, then it's better to 'fit to the model' rather than 'fit the model to us'. We do the former all the time - from banking to electricity provision; we don't go and build our own customized solutions, we just find a way to work with the dominant market standards.

Will cloud computing help the business to treat common, cost-of-doing-business activities as though they are common and a cost of doing business? Or, will some businesses continue to spend vast sums customizing that which really makes no difference?

Could the cloud help the business re-align itself with the market?

Reprinted from LEF site

Friday, February 25, 2011

VMware as an acquisition target?

Back in 2009, I proposed that VMware (or more importantly its master EMC) would eventually divest itself of its virtualisation business. The reason for my thinking was as follows :-
  1. Whilst the majority of VMware's revenue was based upon its virtualisation technology, this was an area that was ripe for disruption through two points of attack - open source systems such as KVM and the formation of marketplaces offering utility based virtual infrastructure. The latter almost certainly requires open source reference models to avoid issues around loss of strategic control and when combined with aggressive service competition around price, this doesn't leave much room for a license based proprietary technology.

  2. A dominant position in the enterprise is no guarantee for future success - see Novell Netware and the IPX/SPX vs TCP/IP battle. Critical in such battles is the development of wide public ecosystems and inherently open source has a natural advantage. However, given the revenue position of VMware, it could not afford to undertake this route.

  3. The obvious route for VMware would be to develop a platform play, most likely an open source route with an extensive range of value add services - from assurance to management. The current business would be used to fund the development of this approach until such time as the company could split into two parts - virtualisation & platform.

  4. Given the likely growth of private clouds as a transitional model in the development of the cloud industry, VMware would position the virtualisation business in this space for a high value sale, benefiting from its strength in the enterprise. This is despite it being unlikely that VMware would become the defacto public standard, that the technology was likely to be disrupted, that hybrid clouds are a transitional strategy superseded by formation of competitive markets and that many "private clouds" would be little more than rebranded virtual data centres.

  5. During this time, there would be signals of confusion over the VMware strategy precisely because it would be using a time limited cash cow to fund a new venture whilst preparing to jettison the cash cow prior to disruption.
Given the confusion over cloud, the often central (but ultimately misconceived) role that virtualisation is given in the industry, the generally disruptive effects of this change and the wealth of many competitors then in my opinion with luck, timing and good judgement a buyer could be found.

So, in my opinion:
  • For VMware, it would mean creating a strong platform business funded by its current revenue stream before jettisoning the virtualisation business at high value prior to its disruption.

  • For the buyer, it would mean ... whoops ... well that's capitalism for you. Next time, pay more attention.
Of course, this is just my opinion but I haven't changed my view over the years. I'm expecting to see an increasingly clear division within VMware between platform and virtualisation in the next year or so.

I'm hence curious to know what others think? Do you believe that VMware would sell its virtualisation business?

Sunday, January 16, 2011

A full and workable definition for cloud?

Definition provided courtesy of Andrey Markov and the texts of many prominent cloud figures.

P.S. Before taking this seriously, please read the note at the end.

=== Start Text ===

Cloud data centers have lower spend per data centers. The authors lay out peaks and reassigned according to recognize that facilitate incremental improvement and four deployment models. Any CIO will use software to take full advantage and ultimately overrides initial concerns (e.g., mission, security holes, high compute task implementation).

When computing offer enough economic advantages will come to achieve higher level of use and smooth out peaks and elastically provisioned, in terms of IT being service oriented (as the paper "The Economics of service models") and reliability will undoubtedly be able to be increased significantly, thereby reducing operational costs caused by heterogeneous thin client interface or thick client platforms (e.g. networks) and can be recognized, and cost of adoption despite initial reservations, resulting in a private cloud with different physical characteristics and reduces cost of the same subject.

"It's the entire industry group and reduces cost advantage include operating systems, storage, applications, and applications".

The basis for that underlie cloud computing? There is likely to a third party and possibly application architecture design pattern learning present themselves as examples. Their point is nothing new (application capabilities, fragility, security spend per data center). The cloud bursting for a significant input to upgrade. We put up as needed automatically without doubt.

Beyond the consumer is in part, because of power management is composed of the TCO (per server, based on the cloud with fragility, security and compliance considerations). It may be unlimited and reliability will adopt technologies despite technical concerns is to vastly increase the rapid innovation we've grown accustomed to. To this is remarkable.

In concluding the authors then ventures into what types of disruptions, as a much easier to use and supports a cloud computing initiative, or service models (and services) that of scale available over time. Their point is available to yourself to be its article; nevertheless, the consumer does cloud computing in the TCO factor compared to the original UC Berkeley Cloud Platform as noted, this question, the economic advantage, organizations will gradually fade, thus resulting in terms of service (e.g. host firewalls).

Deployment Models: Private cloud. The basis for the deployed applications, and accessed through the possible exception of the provided resources include operating systems or acquired applications. Programming languages are reported providing transparency for use of 1000 hours, more compute-intensive tasks will result from (titled "Data center is identified, negotiate extremely low rates than that"). I feel this white paper will be summed up as well. Instead the provider’s applications running on to this white paper is nothing but application hosting environment configurations. Cloud Software as a private variant runs a single document and helps shape the same subject:

"It's the implications of configurable computing economics".

Essential Characteristics: On-demand self-service. A copy of business units, or live with cloud infrastructure is operated solely for 1000 servers for a thin or getting ready to be very large volumes of the ratio of public cloud computing. I would be recognized, and may be able to upgrade. We put up with minimal management is composed of scale advantage.

Supply-side savings: Cloud computing capabilities, such as server data center buildouts, the scale of these cost = increased use. Elasticity of the cloud infrastructure is made available over operating systems, storage, processing, memory, network bandwidth, and lay out peaks and consumer of the underlying economics that economic advantage, organizations and concerns is going to deploy onto the consumer with cloud infrastructure.

In end, the cloud computing does not manage or getting ready to which cloud infrastructure is.

=== End Text ===

The real shame of this, is it actually makes more sense than some of the stuff I get to read. Thanks to @b6n for the suggestion of running a set of popular cloud posts through a Markov Text Generator.

Wednesday, January 05, 2011

My top 10 influential thinkers in cloud ...

I've never been a great fan of top ten lists, however since it appears fashionable I thought I'd have a go. So, here are my top 10 most influential thinkers to the cloud in order of priority:

  1. Douglas Parkhill: For predicting the entire field and writing the exceptional book "The Challenge of the Computer Utility" (1966)
  2. Joseph Schumpeter: For providing a basic economic framework which explains why cloud computing should enable further innovation through creative destruction (1942)
  3. Herbert Simon: For providing a basic economic framework which explains why cloud computing should accelerate further innovation through componentisation (1973)
  4. William Stanley Jevons: For outlining why cloud computing won't reduce overall IT expenditure (1865)
  5. Leigh Van Valen: For providing in "a new evolutionary law" a framework to explain why you wouldn't have choice over cloud computing (1973)
  6. Tim O'Reilly: For signposting the cloud computing future with the concept of infoware and highlighting the role of the internet and open source in this (1999)
  7. Everett Rogers: For postulating that diffusion and maturation of a technological innovation results in increased information about the technology and therefore reduces uncertainty about the change. A key cornerstone of the idea of commoditisation (1981)
  8. Paul Strassmann: For demonstrating that there was no correlation between IT spending and business value, hence showing that not all IT is the same and that some was little more than a cost of doing business i.e. it had become more of a commodity (1990s)
  9. Nick Carr: For showing that ubiquity was the key to diminishing strategic value in business and providing the crucial link to explain commoditisation (2003)
  10. John McCarthy: For being the first person to publicly state the idea of utility computing (1961)

Friday, November 19, 2010

All in a word.

In my previous post, I provided a more fully fledged version of the lifecycle curve that I use to discuss how activities change. I've spoken about this for many years but I thought I spend a little time focusing on a few nuances.

Today, I'll talk about the *aaS misconception - a pet hate of mine. The figure below shows the evolution of infrastructure through different stages. [The stages are outlined in the previous post]

Figure 1 - Lifecycle (click on image for higher resolution)


I'll note that service bureau's started back in the 1960s and we have a rich history of hosting companies which date well before the current "cloud" phenomenon. This causes a great deal of confusion over who and who isn't providing cloud.

The problem is the use of the *aaS terms such as Infrastructure as a Service. Infrastructure clouds aren't just about Infrastructure as a Service, they're about Infrastructure as a Utility Service.

Much of the confusion has been caused by the great renaming of utility computing to cloud, which is why I'm fairly consistent on the need to return to Parkhill's view of the world (Challenge of the Computer Utility, 1966).

Cloud exists because infrastructure has become ubiquitous and well defined enough to support the volume operations needed for provision of a commodity through utility services. The commodity part of the equation is vital to understanding what is happening and it provides the distinction between a VDC (virtual data centre) and cloud environments.

If you're building an infrastructure cloud (whether public or private) then I'll assume you've got multi-tenancy, APIs for creating instances, utility billing and you are probably using some form of virtualisation. Now, if this is the case then you're part of the way there, so go check out your data centre.

IF :-
  • your data centre is full of racks or containers each with volumes of highly commoditised servers
  • you've stripped out almost all physical redundancy because frankly it's too expensive and only exists because of legacy architectural principles due to the high MTTR for replacement of equipment
  • you're working on the principle of volume operations and provision of standardised "good enough" components with defined sizes of virtual servers
  • the environment is heavily automated
  • you're working hard to drive even greater standardisation and cost efficiencies
  • you don't know where applications are running in your data centre and you don't care.
  • you don't care if a single server dies

... then you're treating infrastructure like a commodity and you're running a cloud.

The economies of scale you can make with your cloud will vary according to size, this is something you've come to accept. But when dealing with scale you should be looking at :-
  • operating not on the basis of servers but of racks or containers i.e. when enough of a rack is dead you pull it out and replace it with a new one
  • your TCO (incl hardware/software/people/power/building ...) for providing a standard virtual server is probably somewhere between $200 - $400 per annum and you're trying to make it less.
Obviously, you might make compromises for reasons of short term educational barriers (i.e. to encourage adoption). Examples include: you might want the ability to know where an application is running or to move an application from one server to another or you might even have a highly resilient section to cope with many legacy systems that have developed with old architectural principles such as Scale-up and N+1. Whilst these are valuable short term measures and there will be many niche markets carved out based upon such capabilities, they incur costs and ultimately aren't needed.

Cost and variability are what you want to drive out of the system ... that's the whole point about a utility. Anyway, rant over until next week.

Tuesday, September 21, 2010

A run on your cloud?

When I use a bank, I'm fully aware that the statement I receive is just a set of digits outlining an agreement of how much money I have or owe. In the case of savings, this doesn't mean the bank has my money in a vault somewhere as in all likelihood it's been lent out or used elsewhere. The system works because a certain amount of reserve is kept in order to cover financial transactions and an assumption is made that most of my money will stay put.

Of course, as soon as large numbers of people try to get their money out, it causes a run on the bank and we discover just how little the reserves are. Fortunately, in the UK we have an FSA scheme to guarantee a minimum amount that will be returned.

So, what's this got to do with cloud? Well, cloud (as with banking) works on a utility model, though in the case of banking we get paid on both the amount we consume and provide (i.e interest) and in the cloud world we normally only have the option to consume.

In the case of infrastructure service providers, there are no standard units (i.e. there is no common cloud currency) but instead each provider offers it own range of units. Hence if I rent a thousand computer resource units, those units are defined by that provider as offering a certain amount of storage and CPU for a given level of quality at specified rate (often an hourly fee).

As with any utility there is no guarantee that when I want more, the provider is willing to offer this or has the capacity to do so. This is why the claims of infinite availability are no more than an illusion.

However, hidden in the depths of this is a problem with transparency which could cause a run on your cloud in much the same way that Credit Default Swaps hit many financial institutions as debt exceeded our capacity to service it.

When I rent a compute resource unit from a provider, I'm working on the assumption that what I'm getting is that compute resource unit and not some part of it. For example, if I'm renting on an hourly basis a 1Ghz core with 100Gb storage and 2Gb memory - I'm expecting exactly that.

However, I might not use the whole of this compute resource. This offers the service provider, if they were inclined, an opportunity to sell the excess to another user. In this way, a service provider running on a utility basis could be actively selling 200 of their self defined compute units to customers whilst it only has the capacity to provide for 100 of those units when fully used. This is quaintly given terms like improving utilisation or overbooking or oversubscription but fundamentally it's all about maximising service provider margin.

The problem occurs when everyone tries to use their compute resources fully with an overbooked provider, just like everyone trying to get their money out of a bank. The provider is unable to meet its obligations and partially collapses. The likely effect will be compute units being vastly below their specification or some units which have been sold are thrown off the service to make up for the shortfall (i.e. customers are bumped).

It's worth remembering that a key part of cloud computing is a componentisation effect which is likely to lead to massively increased usage of computer infrastructure in ever more ephemeral infrastructures and as a result our dependency on this commodity provision will increase. It's all worth remembering that black swan events, like bank runs do occur.

If one overbooked provider collapses, then this is likely to create increased strain on other providers as users seek alternative sources of computer resource. Due to such an event and unexpected demand, this might lead to a temporary condition where some providers are not able to hand out additional capacity (i.e. new compute units) - the banking equivalent of closing the doors or localised brown-outs in the electricity industry.

However, people being people will tend to maximise the use of what they already have. Hence, if I'm renting 100 units with one provider who is collapsing, 100 units with another who isn't and a situation where many providers are closing their doors temporarily, then I'll tend to double up the workload where possible on my fully working 100 units (i.e where I believe I have spare capacity).

Unfortunately, I won't be the only one doing this and if that provider has overbooked then it'll collapse to some degree. The net effect is a potential cascade failure.

Now, this failure would not be the result of poor utility planning but instead the overbooking and hence overselling of capacity which does not exist, in much the same way that debt was sold beyond our capacity to service it. The providers have no way of predicting black swan events, nor can they estimate the uncertainty with user consumption (users, however, are more capable of predicting there own likely demands).

There are several solutions to this, however all require clear transparency on the level of overbooking. In the case of Amazon, Werner has made a clear statement that they don't overbook and sell your unused capacity i.e. you get exactly what you paid for.

Rackspace also states that they offer guaranteed and reserved levels of CPU, RAM and Storage with no over subscription (i.e. overbooking).

In this case of VMWare's vCloud Director, then according to James Watters they provide a mechanism for buying a hard reservation from a provider (i.e. a defined unit), with any over commitment being done by the user and under their control.

When it comes to choosing an infrastructure cloud provider, I can only recommend that you first start by asking them what units of compute resource they sell? Then afterwords, ask them whether you actually get that unit or merely a capacity for such depending upon what others are doing? In short, does a compute unit of 1Ghz core with 100Gb storage and 2Gb memory actually mean that or could it mean a lot less?

It's worth knowing exactly what you're getting for your buck.

Wednesday, August 18, 2010

Arguably, the best cloud conference in the world?

For those of you who missed the OSCON Cloud Summit, I've put together a list of the videos and speakers. Obviously this doesn't recreate the event, which was an absolute blast, but at least it'll give you a flavour of what was missed.

Welcome to Cloud Summit [Video 14:28]
Very light introduction into cloud computing with an introduction to the speakers and the conference itself. This section is only really relevant for laying out the conference, so can easily be skipped.
With John Willis (@botchagalupe) of opscode and myself (@swardley) of Leading Edge Forum.

Scene Setting
In these opening sessions we looked at some of the practical issues that cloud creates.

Is the Enterprise Ready for the Cloud? [Video 16:39]
This session examines the challenges that face enterprises in adopting cloud computing. Is it just a technology problem or are there management considerations? Are enterprises adopted cloud, is the cloud ready for them and are they ready for it?
With Mark Masterson (@mastermark) of CSC.

Security, Identity – Back to the Drawing Board? [Video 25:12]
Is much of the cloud security debate simply FUD or are there some real consequences of this change?
With Subra Kumaraswamy(@subrak) of Ebay.

Cloudy Operations [Video 22:10]
In the cloud world new paradigms and memes are appearing :- the rise of the “DevOps”, “Infrastructure == Code” and “Design for Failure”. Given that cloud is fundamentally about volume operations of a commoditized activity, operations become a key battleground for competitive efficiency. Automation and orchestration appear key areas for the future development of the cloud. We review current thinking and who is leading this change.
With John Willis (@botchagalupe) of opscode.

The Cloud Myths, Schemes and Dirty Little Secrets [Video 17:38]
The cloud is surrounded by many claims but how many of these stand up to scrutiny. How many are based on fact or are simply wishful thinking? Is cloud computing green, will it save you money, will it lead to faster rates of innovation? We explore this subject and look at the dirty little secrets that no-one wants to tell you.
With Patrick Kerpan (@pjktech) of CohesiveFT.

Curing Addiction is Easier [Video 18:41]
Since Douglas Parkhill first introduced us to the idea of competitive markets of compute utilities back in the 1960s, the question has always been when would this occur? However, is a competitive marketplace in the interests of everyone and do providers want easy switching? We examine the issue of standards and portability in the cloud.
With Stephen O’Grady (@sogrady) of Redmonk.

Future Setting
In this section we heard from leading visionaries on the trends they see occurring in the cloud and the connection and relationships to other changes in our industry.

The Future of the Cloud [Video 29:00]
Cloud seems to be happening now but where is it going and where are we heading?
With J.P. Rangaswami (@jobsworth) of BT.

Cloud, E2.0 – Joining the Dots [Video 30:04]
Is cloud just an isolated phenomenon, or is it connected to many of the other changes in our industries.
With Dion Hinchcliffe (@dhinchcliffe) of Dachis.

The Questions
The next section was a Trial by Jury where we examined some of the key questions around cloud and open source.

What We Need are Standards in the Cloud [Video 45:17]
We put this question to the test, with prosecution Benjamin Black (@b6n) of FastIP, defence Sam Johnston (@samj) of Google and trial by a Jury of John Willis, Mark Masterson, Patrick Kerpan & Stephen O’Grady

Are Open APIs Enough to Prevent Lock-in? [Video 43:21]
We put this question to the test, with prosecution James Duncan (@jamesaduncan) of Joyent, defence George Reese (@georgereese) of Enstratus and trial by a Jury of John Willis, Mark Masterson, Patrick Kerpan & Stephen O’Grady

The Debates
Following the introductory sessions, the conference focused on two major debates. The first of these covered the “cloud computing and open source question”. To introduce the subject and the panelists, there were a number of short talks before the panel debates the impact of open source to cloud and vice versa.

The Journey So Far [Video 10:59]
An overview of how “cloud” has changed in the last five years.
With James Urquhart (@jamesurquhart) of CISCO.

Cloud and Open Source – A Natural Fit or Mortal Enemies? [Video 8:44]
Does open source matter in the cloud? Are they complimentary or antagonistic?
With Marten Mickos (@martenmickos) of Eucalyptus.

Cloudy Futures? The Role of Open Source in Creating Competitive Markets [Video 8:43]
How will open source help create competitive markets? Do “bits” have value in the future and will there be a place for proprietary technology?
With Rick Clark (@dendrobates) of OpenStack.

The Future of Open Source [Video 9:34]
What will cloud mean to open source development and to linux distributions. Will anyone care about the distro anymore?
With Neil Levine (@neilwlevine) of Canonical.

The Debate – Open Source and the Cloud
 [Video 36:24]
Our panel of experts examined the relationship between open source and cloud computing.
With Rick Clark, Neil Levine, Marten Mickos & James Urquhart

The Future Panel followed the same format with first an introduction to the experts who will debate where cloud is going to take us.

The Government and Cloud [Video 10:27]
The role of cloud computing in government IT – an introduction to the large G-Cloud and App Store project under way in the UK; what the UK public sector hopes to gain from a cloud approach, an overview of the proposed technical architecture, and how to deliver the benefits of cloud while still meeting government’s stringent security requirements.
With Kate Craig-Wood (@memset_kate) of Memset.

Infoware + 10 Years [Video 10:38]
Ten years after Tim created the term infoware, how have things turned out and what is the cloud’s role in this?
With Tim O'Reilly (@timoreilly) of O'Reilly Media.

The Debate – A Cloudy Future or Can We See Trends? [Video 50:12]
The panel of experts examine what’s next for cloud computing, what trends can they forsee.
With Kate Craig-Wood, Dion Hinchcliffe, Tim O’Reilly & JP Rangaswami

So, why "arguably the best cloud conference in the world?"

As a general conference on cloud, then the standard and quality of the speakers was outstanding. The speakers made the conference, they gave their time freely and were selected from a wide group of opinion leaders in this space. There was no vendor pitches, no paid for conference speaking slots and hence the discussion was frank and open. The audience themselves responded marvelously with a range of demanding questions.

It is almost impossible to pick a best talk from the conference because they were all great talks. There are real gems of insight to be found in each and every one and each could easily be the keynote for most conferences. In my opinion, if there is a TED of cloud, then this was it.

Overall, the blend of speakers and audience made it the best cloud conference that I've ever attended (and I've been to 50+). This also made my job as a moderator simple.

I'm very grateful to have been part of this and so my thanks goes to the speakers, the audience, the A/V crew who made life so easy and also Edd Dumbill (@edd), Allison Randal (@allisonrandal), Gina Blaber (@ginablaber) and Shirley Bailes (@shirleybailes) for making it happen.

Finally, huge thanks to Edd and Allison for letting me give a version of my Situation Normal, Everything Must Change talk covering cloud, innovation, commoditisation and my work at LEF.